This is the benchmark post: named passages, raw numbers, and why the gap exists. If you only want a buying decision with a feature table and a verdict, jump to our /vs/writehuman comparison page. What follows is the evidence under that verdict.
Why test humanizers out of distribution?
An out-of-distribution test uses passages on domains neither tool was tuned on, which is the opposite of a vendor demo. Most comparisons run a tidy ChatGPT intro on a generic topic, the kind every model has seen ten thousand times: both tools pass and the result tells you nothing about your actual essay.
We picked these passages precisely because they sit outside the easy center of the training distribution. The point was generalisation, not best case. A humanizer that only works on common topics is a humanizer that fails on the upper-division paper you are actually trying to submit.
So we chose five academic passages across five unrelated fields, ran each through both tools, and recorded what the detectors returned. No cherry-picking, no re-rolls to find a flattering number. The first output each tool produced is the output we scored.
Which five passages did we test?
The five test passages cover psychology, linguistics, art history, environmental economics and education theory, each a 150-to-300-word AI-generated academic paragraph. A benchmark you cannot inspect is a marketing claim wearing a lab coat, so here they are by name:
- psy01, a psychology passage on cognitive load theory and working memory limits.
- lin01, a linguistics passage on code-switching among bilingual speakers.
- art01, an art history passage on chiaroscuro technique in Baroque painting.
- env01, an environmental economics passage on carbon-pricing externalities.
- edu01, an education theory passage on constructivist pedagogy.
These domains were deliberate. Psychology, linguistics, and art history cover the humanities and social-science register where most upper-division essays live. Environmental economics and education theory are denser, more jargon-heavy, and structurally formal, which is exactly the kind of prose that triggers detectors hardest. If a humanizer survives env01 and edu01, the lighter cases take care of themselves.
How was each rewrite scored?
Scoring was percent AI-flagged, so lower is better and 0% means the detector read the rewrite as fully human. Each AI passage went through HumanizeMyAI and through WriteHuman Enhanced, the tool's strongest available mode, and both outputs went into QuillBot's AI detector for the headline number.
QuillBot carries the headline because it is widely used by students and gives a clean percentage. We did not stop at one detector. Each output also went through GPTZero, Turnitin AI 2026, Originality AI, and Copyleaks. A single detector can be fooled or can misfire; agreement across several is what tells you whether a result is real or noise. One number is an anecdote. Five detectors pointing the same way is a finding.
A spike into the double digits means the detector caught something the rewrite left behind.
What did the head-to-head benchmark return?
HumanizeMyAI scored 0% AI on all 5 passages across QuillBot. WriteHuman scored 0% on 3 of the 5, but spiked to 61% on env01 and 34% on edu01. On the cover that reads 5/5 vs 3/5, and the cover reports exactly that.
The cross-detector pass mirrored the headline. On the other four detectors, WriteHuman degraded on the same two passages, env01 and edu01, while holding up on psy01, lin01, and art01. That consistency matters. It means the spike was not a QuillBot fluke. The two dense, formal domains carried something the Enhanced rewrite could not flatten, and every detector we tried noticed.
For HumanizeMyAI, the five-passage result fits the wider record. In the owner-verified run of 31 August 2026, our output read 0% on GPTZero, Copyleaks and QuillBot, 0-3% on ZeroGPT, and drew Human verdicts from Turnitin and Originality AI, a 0.3% mean across the four detectors that print a number.
Why did WriteHuman spike on env01 and edu01?
WriteHuman spiked on env01 and edu01 because it is a paraphraser-class tool: at its core it rewrites by swapping words and resequencing clauses while keeping the original skeleton mostly intact. That works well when the source prose is loosely structured, which is why psy01, lin01, and art01 passed cleanly. Dense formal prose gives it no room.
env01 and edu01 are different. Carbon-pricing economics and constructivist pedagogy lean on rigid scaffolding: defined terms that cannot be swapped without breaking meaning, parallel three-part lists, and scholar-attribution chains. A paraphraser hits a wall there. It cannot substitute "marginal abatement cost" for a casual synonym, so the formal lexical fingerprint survives the rewrite, and detectors lock onto exactly that fingerprint. The skeleton the model started with is the skeleton the detector still sees.
HumanizeMyAI did not paraphrase those passages. It restructured them, lowering register and rebuilding the rhythm so the underlying statistical signature changed, not just the surface vocabulary. That is the whole reason the corpus matters: it gives the rewrite real human paragraph patterns to land on instead of a synonym table.
Where does WriteHuman hold up?
On the three passages it passed, WriteHuman passed cleanly. It supports multiple languages, which matters if you write outside English, and its Enhanced mode rewrites more aggressively than the base setting. It is an established commercial product, which is why it serves as the reference baseline in this benchmark.
The catch is consistency. Three of five passing is good until the two that fail are the dense academic passages you most needed to pass. A tool that works on your blog intro but spikes on your thesis chapter is a tool you cannot fully rely on for high-stakes work, and that is the trade-off the benchmark makes visible.
What does run-to-run variance mean for your essay?
Run-to-run variance means the same tool, in the same mode, behaves differently depending on the topic you feed it. Paraphraser-class tuning specializes on a calibration sweet spot that does not cover every domain: three passages landed inside it, and two dense formal ones fell outside and spiked.
The practical reading is about risk. Loose, reflective prose sits inside a paraphraser's comfort zone, which is where WriteHuman passed. A technical, quantitative, or heavily-cited paper brings the env01 and edu01 dynamics with it, and that is precisely where a paraphraser's variance turns into a failed submission.
HumanizeMyAI's 0% across all five is the other side of that coin: lower variance across domains, because corpus-grounded restructuring does not depend on the source topic sitting inside a narrow tuning band. The benchmark cannot tell you which assignment you have. It can tell you which tool behaves more predictably when the topic gets hard, and predictability is what high-stakes work actually needs.
How much does WriteHuman cost?
WriteHuman offers a limited free trial capped at 200 words, then Premium at $24/mo for 10,000 words per month, and Premium+ at $43/mo for 50,000 words per month. Price shapes how you can actually test a tool, and the /vs/writehuman decision page carries the full comparison.
Two practical things fall out of that. First, the 200-word free trial is too small to run a real essay paragraph, so you largely have to pay before you can test the tool on your own writing. Second, the monthly word ceiling is real: at the $24/mo tier, a single 5,000-word research paper eats half your month in one sitting.
By contrast, HumanizeMyAI's free tier is 4 runs at 250 words each on an account that costs nothing to open, so you can put a genuine chunk of your own essay through before deciding anything. The difference that matters is not price. It is that you can verify us on your own writing first. For the head-to-head pricing math and the verdict rating, that all lives on the /vs/writehuman page; this post stays on the benchmark.
Can QuillBot's detector catch QuillBot's own humanizer?
Yes. QuillBot ships both a humanizer and an AI detector, so its detector can be asked to grade its own paraphraser, and QuillBot Humanizer's own output returns approximately 95% AI on it (owner re-test, May 15, 2026). That number is why a 0% result is more meaningful than it looks.
The vendor's classifier catches the vendor's own paraphraser at roughly 95% AI, and on that same classifier HumanizeMyAI returns 0%. That asymmetry is the tell. It is not synonym-swap tuning that happens to dodge one detector. It is an architectural difference: a paraphraser leaves a paraphraser signature even its own maker can detect, and corpus-grounded restructuring does not. For the full teardown of QuillBot as both humanizer and detector, see our QuillBot Humanizer review.
How do you pick the right humanizer for your case?
Picking the right humanizer comes down to running your own hardest passage through the candidates and scoring it on more than one detector, then checking the word caps against your monthly volume. Here is a 6-step way to reach that decision without trusting anyone's marketing:
- Name your hardest passage. Take the single densest, most formal paragraph in your document, the env01-or-edu01 of your own work, not the easy intro.
- Check your free runway. With HumanizeMyAI a free account carries 4 passages at 250 words, and no card is asked for. With WriteHuman the free trial is 200 words, so plan to pay before you can fully test it.
- Run that hard passage through both, Enhanced mode on WriteHuman, default on HumanizeMyAI.
- Score across detectors, not one. Put each output through QuillBot plus GPTZero and one more. Agreement is the signal; a lone pass is not.
- Do the cap math for your volume. If you write 5,000-plus words a month, weigh WriteHuman's 10,000-word Premium ceiling against how much headroom you actually need.
- Pick on your own numbers. Whatever scores lower across detectors on your hardest passage wins. If you want the full feature-by-feature comparison and the verdict rating before you commit, that is exactly what the /vs/writehuman decision page is for.
If Turnitin's Aug 2025 layered classifier is the gate you are actually worried about, the academic workflow is covered in our Turnitin AI bypass guide. For the full ranked field of nine humanizers across all six detectors, see our best AI humanizer 2026 listicle. And the pattern scanner behind this benchmark also runs on our AI detector page, alongside the trained model that produces the verdict there, so you can paste any text and check it yourself.
By Fırat Mıhcı (ResearchGate). $0 affiliate stake.