For two years the advice has been the same: to spot AI writing, watch for delve, intricate, underscore, showcase, nuanced. A whole cottage industry of detectors and classroom rules rests on that vocabulary. But a word can only be a tell if it stays put: if it lives in machine text and not in human text. So we asked the obvious follow-up question to our own earlier word study: two years on, where did those words actually go? We tracked them across 18,989 scientific abstracts spanning eight years and three fields, and the answer is the quiet alarm in this paper. The AI vocabulary did not stay in AI text. It leaked into the human record, and you can watch it happen.
The clean experiment hiding in arXiv
To prove a word is diffusing rather than just trending, you need a place where it has no business showing up. That place is mathematics.
We pulled abstracts from three arXiv fields, every year from 2019 to 2026: cs.CL (computational linguistics: a field whose researchers were early, heavy adopters of LLMs for writing), q-bio.NC (neuroscience: a moderate-adoption control), and math.PR (probability theory: the cleanest control we could find). Then we isolated a pure-stylistic subset of the AI register: words like delve, intricate, underscore, showcase and nuanced that carry no field-specific meaning. A probability paper has every topical reason to say theorem, bound or stochastic, and no topical reason whatsoever to say delve or tapestry. If those purely decorative words start rising in a math abstract, topic drift can't explain it. Only style can.
That design is the whole point. Anyone can show that delve got more common somewhere. The hard claim (the one worth a study) is that it got more common in writing that had no subject-matter reason to reach for it.
What the eight-year curve shows
From 2019 through 2022, the picture is calm and identical everywhere. In all three fields, these pure-stylistic words appear at roughly one to two per 10,000 words: a flat, boring baseline, year after year. Linguistics, neuroscience and mathematics are indistinguishable. There is no pre-existing trend for ChatGPT to merely accelerate.
Then late 2022 happens, and the fields split apart hard.
- Computational linguistics climbs from 1.9 per 10,000 words in 2022 to 14.0 in 2024: roughly a sevenfold jump in two years, and the 95% confidence intervals for the two years don't even overlap. This is not noise.
- Neuroscience rises too, about fivefold over the same window: real, but milder.
- Mathematics goes from 0.9 to 1.3 per 10,000 words. Its confidence intervals overlap. In plain terms, the math control barely moves.
Put another way: by 2024, 17.5% of computational-linguistics abstracts contained at least one of these purely stylistic AI-register words, up from 2.6% in 2022. About one in six. In probability theory, the same words stayed where they had always been.
The dose-response is the tell
Notice the ordering. The surge is largest in the field that writes with LLMs the most, smaller in the field that uses them moderately, and absent in the field that has the least cultural reason to touch them. Linguistics ≫ neuroscience ≫ mathematics, and that ranking lines up exactly with how heavily each community adopted AI writing tools.
That gradient is what turns a correlation into an argument. A generic "everyone writes a bit more formally now" effect would lift all three fields together. An arXiv-wide formatting change, a new template, a shift in what gets submitted: any site-wide artifact would also move mathematics. None of them did. The signal scales with LLM exposure and vanishes without it, which is the fingerprint of stylistic diffusion: the machine's habits seeping outward into the people who use it, concentrated precisely where those people lean on it hardest.
Rise, then pullback
The story has a second act that the headline often misses. The most iconic words don't just rise: they peak in 2023–24 and then partly recede in 2025–26. Call it the delve backlash: once "delve" became a punchline and an accusation, careful writers started scrubbing it back out, and the most flagged words began to retreat from their peak. (Our 2026 figure is a partial year, running through roughly June, so read the final point as provisional rather than a finished trend.)
This rise-and-pullback matters because it shows the human record is not a one-way ratchet. The vocabulary moved in, got noticed, and is now being self-consciously edited back. But "partly recede" is not "return to baseline." The 2025–26 levels sit well above the calm 2019–2022 floor. The words came in, the backlash trimmed the loudest of them, and a residue stayed behind in human-edited prose.
The honest caveat, which is also the point
Here is the thing we will not hide, because it changes what the study can and cannot claim. After late 2022, a large share of these abstracts were almost certainly AI-assisted: drafted, polished or co-written with an LLM and then edited by a human. So this is not a measurement of unaided human beings spontaneously changing how they write. It is a measurement of the AI register entering text that humans edited, approved and put their names on.
For most purposes that distinction would matter enormously. For this one (the contamination of the baseline), the mechanism is beside the point. Whether a researcher typed delve unprompted or accepted it from a draft and let it stand, the word is now in the published human record under a human author. A detector reading that abstract has no way to know which happened, and no reason to care: the human corpus it trains and tests against now contains these words either way.
Two more limits, stated plainly. This is the scientific-abstract genre only (terse, formal, conventionalized prose), and abstracts may well move faster or slower than blog posts, emails or essays. And we did not discover the core rise: Liang et al. (2024) and Kobak et al. (2024) already documented that AI-register words surged in scientific writing after ChatGPT. What this study adds is the mathematics control that rules out topic drift, the field dose-response that ties the surge to LLM adoption, and the 2025–26 pullback that the earlier snapshots were too early to see.
Why a humanizer lab studied a leaking word list
We build language tools. HumanizeMyAI is trained on a published corpus of 2,590 real student essays, and the skill underneath it (reading human writing carefully, at scale, and noticing what is genuinely human about it) is exactly the skill this question demands. We had a direct stake here, too: our own predecessor study found that the famous AI vocabulary was a weak, fast-aging tell to begin with. This paper is the sequel that explains why it ages. The words don't just go out of fashion in the model; they bleed out of the model and into the very human text that any word-based detector treats as its clean reference.
That is the same instinct behind our free AI detector: read the text, not the legend. And it rhymes with our other findings: the sibling study There Is No Aesthetic Surprisal Arc showed that the rhythm of surprise doesn't betray AI either, and that the durable signals are duller and structural. A word list, it turns out, is the most fragile signal of all, because the words refuse to stay on one side of the line.
What this means for spotting AI text
The practical takeaway is uncomfortable for anyone selling a vocabulary checklist. If delve, intricate and underscore are now appearing in human-authored mathematics abstracts (fields with no topical reason to use them), then a detector that scores text by counting AI-register words is increasingly going to false-positive on genuine human writing. The human lexical baseline it relies on has been contaminated by the very phenomenon it claims to catch. Every year, the gap between "human vocabulary" and "AI vocabulary" that these tools depend on gets narrower, and a careful human writer who happens to like the word nuanced looks more and more like a machine.
This is why durable detection cannot live in a word list. Vocabulary is a moving target that both sides now share; the more stable signals are structural: how evenly information is spread, the rhythm of long and short sentences, the architecture of an argument rather than its diction. For the companion story on how a real-corpus approach reads these human-versus-machine textures from the product side, the 2,590-essay corpus deep dive is the read, and the rest of our work lives on the research hub. The honest position is the cautious one: a paragraph full of delve tells you less every year, and might just be a 2024 human who picked up the habit.
Open data, open code
This is not a black box. The full method, every per-field and per-year figure, the complete pure-stylistic word set, and a frank limitations section are written up in the preprint on ResearchGate (DOI 10.13140/RG.2.2.14890.38085), and the complete analysis code with a one-command reproduction of every number above is openly available on GitHub. The abstracts come from arXiv's public metadata, so anyone can pull the same 18,989 abstracts, re-run the same counts across the same three fields, and check the result rather than take it on faith. The fingerprint everyone was taught to look for is leaking into the record it was meant to distinguish, and we have left everything in the open so you can watch it happen, and watch it keep shifting with every model that follows.
About the author
Fırat Mıhcı built HumanizeMyAI on a published 2,590-essay corpus; The Fingerprint Is Leaking Into the Record comes from the same close reading of student prose. Publication record at ResearchGate.
Read the full study
Complete methods, every statistic, all figures and a frank limitations section, plus the analysis code and a fully de-identified dataset.