For two years the advice has been the same: to spot AI writing, watch for delve, intricate, underscore, showcase, nuanced. A whole cottage industry of detectors and classroom rules rests on that vocabulary. But a word can only be a tell if it stays put: if it lives in machine text and not in human text. So we asked the obvious follow-up question to our own earlier word study: two years on, where did those words actually go? We tracked them across 18,989 scientific abstracts spanning eight years and three fields, and the answer is the quiet alarm in this paper. The AI vocabulary did not stay in AI text. It leaked into the human record, and you can watch it happen.
How do you prove AI words are spreading into human writing?
Proving that a word is diffusing rather than merely trending takes a field where it has no business showing up, and that field is mathematics. The experiment therefore compares three arXiv communities over eight years, with probability theory held as the control.
We pulled abstracts from three arXiv fields, every year from 2019 to 2026: cs.CL (computational linguistics: a field whose researchers were early, heavy adopters of LLMs for writing), q-bio.NC (neuroscience: a moderate-adoption control), and math.PR (probability theory: the cleanest control we could find). Then we isolated a pure-stylistic subset of the AI register: words like delve, intricate, underscore, showcase and nuanced that carry no field-specific meaning. A probability paper has every topical reason to say theorem, bound or stochastic, and no topical reason whatsoever to say delve or tapestry. If those purely decorative words start rising in a math abstract, topic drift can't explain it. Only style can.
That design is the whole point. Anyone can show that delve got more common somewhere. The hard claim (the one worth a study) is that it got more common in writing that had no subject-matter reason to reach for it.
What did the eight-year curve show?
The eight-year curve stays flat until late 2022 and then splits by field. From 2019 through 2022, these pure-stylistic words appear at roughly one to two per 10,000 words in all three fields, a boring baseline where linguistics, neuroscience and mathematics are indistinguishable. There is no pre-existing trend for ChatGPT to merely accelerate.
Then late 2022 happens, and the fields split apart hard.
- Computational linguistics climbs from 1.9 per 10,000 words in 2022 to 14.0 in 2024: roughly a sevenfold jump in two years, and the 95% confidence intervals for the two years don't even overlap. This is not noise.
- Neuroscience rises too, about fivefold over the same window: real, but milder.
- Mathematics goes from 0.9 to 1.3 per 10,000 words. Its confidence intervals overlap. In plain terms, the math control barely moves.
Put another way: by 2024, 17.5% of computational-linguistics abstracts contained at least one of these purely stylistic AI-register words, up from 2.6% in 2022. About one in six. In probability theory, the same words stayed where they had always been.
Why is the dose-response pattern the proof?
The dose-response pattern is the proof because the surge scales with exposure: largest in the field that writes with LLMs the most, smaller in the field that uses them moderately, and absent in the field with the least cultural reason to touch them. Linguistics ≫ neuroscience ≫ mathematics, and that ranking lines up exactly with how heavily each community adopted AI writing tools.
That gradient is what turns a correlation into an argument. A generic "everyone writes a bit more formally now" effect would lift all three fields together. An arXiv-wide formatting change, a new template, a shift in what gets submitted: any site-wide artifact would also move mathematics. None of them did. The signal scales with LLM exposure and vanishes without it, which is the fingerprint of stylistic diffusion: the machine's habits seeping outward into the people who use it, concentrated precisely where those people lean on it hardest.
Did the AI words keep rising after 2024?
The most iconic words did not keep rising: they peak in 2023 and 2024, then partly recede across 2025 and 2026. Call it the delve backlash, the second act the headline often misses: once "delve" became a punchline and an accusation, careful writers started scrubbing it back out, and the most flagged words began to retreat from their peak.
(Our 2026 figure is a partial year, running through roughly June, so read the final point as provisional rather than a finished trend.)
This rise-and-pullback matters because it shows the human record is not a one-way ratchet. The vocabulary moved in, got noticed, and is now being self-consciously edited back. But "partly recede" is not "return to baseline." The 2025 and 2026 levels sit well above the calm 2019 to 2022 floor. The words came in, the backlash trimmed the loudest of them, and a residue stayed behind in human-edited prose.
What can this study not claim?
This study cannot claim that unaided human beings spontaneously changed how they write. After late 2022, a large share of these abstracts were almost certainly AI-assisted: drafted, polished or co-written with an LLM and then edited by a human. What it measures is the AI register entering text that humans edited, approved and put their names on.
For most purposes that distinction would matter enormously. For this one (the contamination of the baseline), the mechanism is beside the point. Whether a researcher typed delve unprompted or accepted it from a draft and let it stand, the word is now in the published human record under a human author. A detector reading that abstract has no way to know which happened, and no reason to care: the human corpus it trains and tests against now contains these words either way.
Two more limits, stated plainly. This is the scientific-abstract genre only (terse, formal, conventionalized prose), and abstracts may well move faster or slower than blog posts, emails or essays. And we did not discover the core rise: Liang et al. (2024) and Kobak et al. (2024) already documented that AI-register words surged in scientific writing after ChatGPT. What this study adds is the mathematics control that rules out topic drift, the field dose-response that ties the surge to LLM adoption, and the 2025 to 2026 pullback that the earlier snapshots were too early to see.
Why did a humanizer lab study a leaking word list?
A leaking word list is our own unfinished business: our predecessor study found the famous AI vocabulary to be a weak, fast-aging tell, and this paper is the sequel that explains why it ages. Reading human writing at scale, and noticing what is genuinely human about it, is the same skill behind HumanizeMyAI and its published corpus of 2,590 real student essays.
The words don't just go out of fashion in the model; they bleed out of the model and into the very human text that any word-based detector treats as its clean reference.
That is the same instinct behind our free AI detector: weigh the writing in front of you, never the folklore around it. And it rhymes with our other findings: the sibling study There Is No Aesthetic Surprisal Arc showed that the rhythm of surprise doesn't betray AI either, and that the durable signals are duller and structural. A word list, it turns out, is the most fragile signal of all, because the words refuse to stay on one side of the line.
What does a contaminated baseline mean for AI detectors?
A contaminated baseline means a detector that scores text by counting AI-register words will increasingly false-positive on genuine human writing. If delve, intricate and underscore now appear in human-authored mathematics abstracts, fields with no topical reason to use them, then the human lexical baseline these tools rely on has been contaminated by the very phenomenon they claim to catch.
Every year, the gap between "human vocabulary" and "AI vocabulary" that these tools depend on gets narrower, and a careful human writer who happens to like the word nuanced looks more and more like a machine.
This is why durable detection cannot live in a word list. Vocabulary is a moving target that both sides now share; the more stable signals are structural: how evenly information is spread, the rhythm of long and short sentences, the architecture of an argument rather than its diction. For the product-side view of how a real-corpus approach handles that same material, the 2,590-essay corpus deep dive covers it, and the rest of our work lives on the research hub. The honest position is the cautious one: a paragraph full of delve tells you less every year, and might just be a 2024 human who picked up the habit.
Can you re-run this arXiv analysis yourself?
You can, on the same public abstracts we used. The method, the per-field and per-year figures, the complete pure-stylistic word set and a plain statement of what the design cannot settle appear in the preprint on ResearchGate (DOI 10.13140/RG.2.2.14890.38085), and the complete analysis code with a one-command reproduction of every number above is openly available on GitHub.
The abstracts come from arXiv's public metadata, so anyone can pull the same 18,989 abstracts, re-run the same counts across the same three fields, and check the result rather than take it on faith. The fingerprint everyone was taught to look for is leaking into the record it was meant to distinguish, and we have left everything in the open so you can watch it happen, and watch it keep shifting with every model that follows.
About the author
Fırat Mıhcı built HumanizeMyAI on a published 2,590-essay corpus; The Fingerprint Is Leaking Into the Record comes from the same close reading of student prose. Publication record at ResearchGate.
Where is the full text of "The Fingerprint Is Leaking Into the Record"?
The preprint holds the complete method, every figure and the data behind "The Fingerprint Is Leaking Into the Record".