Most advice on this question hands you a list of words to delete. In our Computational Linguistics Lab we took the popular rules for spotting AI writing, tested each against public data, and published the code whatever the result. Most of the rules lost. What survived is duller than a word list and much harder to fake.
What Actually Makes Writing Read as Human Instead of AI-Generated?
Human writing reads as human because it is uneven in the places a model smooths out. Its predictability rises and stalls from sentence to sentence, its tone follows the writer’s stance, and its word choices include odd, low-probability ones. Raw AI writing stays steadier and settles into a hedged, list-building explainer voice.
Stylometry, the measurement of writing style, offers several places to look. Here is what the Lab found at each:
| Signal | What the Lab measured | What it means for a reader |
|---|---|---|
| Shape of surprise across a text (the “arc”) | 0.552 at separating human from ChatGPT answers | Not a tell at all |
| Level of surprise (perplexity) | 0.80 on raw long-form text, 0.58 after a paraphrase | A tell that a rewrite erases |
| Variance of surprise (burstiness) | 0.80 on raw long-form text, 0.78 after a paraphrase | The most durable of the three |
| Famous “AI words” (delve, tapestry, realm) | Below the counting threshold in 11 million words of late-2022 answers | A passing fashion |
| Hedged explainer words (overall, important, may) | 4.6 to 9.5 times more common in ChatGPT answers | A register that stays put |
| Em-dash | About 5.4 per 1,000 words in Claude text, about 1.1 in GPT text | Separates vendors, not people from machines |
Scores in the first three rows are ROC-AUC: 0.50 is guessing, 1.00 is perfect separation.
No row in that table is a word you can simply delete. The word-level layer lives in our guide to AI words and phrases to avoid; this page stays with what sits underneath it.
Do AI Models Write With a Predictable Rhythm, or Is That a Myth?
AI models do write with a predictable rhythm, but it is a flat evenness, not the rising and falling arc that popular theory describes. In a public benchmark, an automated paraphrase attack left that evenness in place, which makes rhythm the layer where a writer’s own voice counts most.
The arc study found no arc on three public benchmarks, even where the theory had the most room to work: across 30 public-domain books, the slow, arc-like part of the surprise signal held just 0.12 of the spectrum, against 0.30 in short answers.
The paraphrase test matters more. The benchmark includes an attack in which a model automatically rewords machine text, and on long-form text that attack pushed the perplexity score toward chance in every domain checked, from 0.97 to 0.84 on movie reviews and from 0.93 to 0.80 on PubMed abstracts, while burstiness stayed near 0.78. A second, larger scoring model gave the same split. That attack is one machine rewording another, a different thing from a rewrite modelled on real human writing.
Evenness alone is not bad writing: among human student essays graded by people, the slightly more uniform ones scored a little higher. It becomes a tell only when it never breaks, paragraph after paragraph, which is why a suspiciously steady passage tells you more than a merely predictable one.
Has AI-Style Vocabulary Started Leaking Into Human Writing?
AI-style vocabulary has leaked into human writing, and a control field shows the cause is style, not subject matter. In 2022, 2.6% of computational-linguistics abstracts on arXiv carried even one purely decorative AI-register word such as delve or intricate; by 2024 that figure stood at 17.5%.
The Lab’s vocabulary-leak study followed those words through 18,989 abstracts from 2019 to 2026. In computational linguistics the rate rose from 1.9 to 14.0 per 10,000 words, and the uncertainty ranges around the two years do not touch. In probability theory it went from 0.9 to 1.3, with ranges that overlap: no change the data can detect. Mathematicians have no topical reason to write delve, so a rise confined to the field that leans hardest on AI writing tools, with neuroscience in between at about fivefold, points to habits spreading from the tools.
Many of those abstracts were likely AI-assisted and then signed by a human, but a reader cannot tell which. The words peaked in 2023 and 2024, partly receded, and stayed above the old floor, which makes a word list the weakest evidence there is.
Why Does AI Writing Feel Generic Even When It’s Grammatically Perfect?
AI writing feels generic because it lacks specificity: the detail only this writer, on this subject, could supply, such as a named source, a figure from their own notes, or a claim stated without a cushion. Without it, a paragraph could sit in almost any document, and readers register that interchangeability as machine-made however clean the grammar is.
The Lab’s lexical signature study shows what fills that space. In 11 million words of matched human and ChatGPT answers, the model leaned on factors (about 6.7 times the human rate) and help (about 5.7 times) as well as overall, important, including and may. Each is a way of not committing: may softens, factors and including promise a list instead of a choice, and overall wraps up instead of concluding. Robust, often called an AI word, ran the other way at about 0.47 times the human rate.
Illustration written for this page, not study data:
- Generic: “Several factors may influence customer retention, including pricing, support quality and the overall experience.”
- Specific: “We lost our biggest client in March because a support ticket sat unanswered for nine days.”
The first could open any report about any company. The second commits to one cause, one date and one number.
Does Every AI Model Have Its Own Writing Accent?
Every AI model writes with an accent, and the accent belongs to the company that trained it far more than to the model itself. Given the same 102 prompts, five models could be sorted by vendor at 0.96 on style features alone, while naming the exact model reached 50% against a 20% guessing baseline.
The accent study covered two vendors: three OpenAI GPT models, including the older GPT-3.5, and two Anthropic Claude models of different sizes, all writing plain English in June 2026. Google’s Gemini and other vendors’ models were not tested, so the study says nothing about how they sound.
Inside a vendor, the differences mostly blur. Separating OpenAI’s three models scored 55% where guessing gives 33%, and the larger Claude against the smaller one reached an AUC of only 0.65; only GPT-3.5, the oldest, stood out on its own. A lab’s small model and its flagship sound like the same writer.
Claude writes the more punctuation-forward, varied prose; GPT runs smoother, hedges more and uses more passive voice. The em-dash, at about 5.4 per thousand words for Claude against roughly 1.1 for GPT, marks the vendor, not the machine.
Can You Tell Human From AI Writing Without a Detector?
You can sometimes tell human from AI writing by reading alone, but it depends on who reads and what. In a 2023 PNAS study, untrained readers judging AI-written self-presentations were right 50 to 52% of the time. In a 2020 ACL study, trained raters still misjudged GPT-2 excerpts more than 30% of the time.
Those two studies measure different things. Jakesch and colleagues asked 4,600 participants to judge professional bios, dating profiles and hospitality reviews (PNAS, 2023); paying them for accuracy (51.6%) or giving instant feedback (51.2%) barely moved the result. Ippolito and colleagues gave raters a short walkthrough and open-domain GPT-2 text (ACL, 2020), and on the longest excerpts accuracy passed 70%. Current models write far more fluently than GPT-2, so treat that figure as an early benchmark.
Pointing readers at the right signal helps: the GLTR tool, which highlights how predictable each word is, raised untrained readers from 54% to 72% (ACL, 2019). Read for whether predictability stays level paragraph after paragraph, whether every claim arrives hedged, and whether any detail could only come from the writer.
One caution: careful second-language writing can be even for honest reasons. In the Lab’s proficiency study of 3,300 graded learner essays, what changed with skill was accuracy, with errors falling from about 21 to about 4 per 100 words. If you wrote a text yourself and it was flagged, see our guide to documenting your own writing process; how often the tools misfire is covered in our review of detector reliability.
What Is Register, and Why Do AI Models Struggle to Hold It Consistently?
Register is the tone, formality and stance a piece of writing takes, and AI models hold their default register too firmly rather than too loosely. Whatever the topic, raw model output drifts back to the balanced, hedged explainer. A human writer’s register moves with what they believe and who they are addressing.
Our studies explain why that default is so sticky. The explainer voice stayed stable across model generations, because it comes from how models are trained to be helpful and careful, not from any year’s slang. The accent study found the same pattern one level up: a vendor’s training pipeline stamps one house voice on its flagship and its small models alike.
A human register is harder to imitate because it belongs to a person. That is why the benchmark’s automated attack, in which one model rewords another’s output, brought predictability down yet left the evenness barely touched; rewriting modelled on real human prose is another matter. A sure writer says so plainly, an unsure one says why, and a section they care about runs long. Every Lab finding on this page comes from public data with the code on GitHub, all collected on our research hub.
You can see these signals in your own text right now. Paste a paragraph into our free AI detector, which works without an account, and read the result against the table at the top of this page. If a draft began life in a chatbot, the HumanizeMy humanizer rewrites it in a voice that reads like yours; a free account gives you four runs of up to 250 words, one time, with no card. For more writing than that, the monthly plans cost $18 for Basic, $27 for Pro and $48 for Ultra.
Fırat Mıhcı leads the HumanizeMy Computational Linguistics Lab; his preprints are listed on ResearchGate. This page is revised whenever an underlying study is re-run or a cited paper changes. The site carries no affiliate links: the one product we sell is our own.