A Computational Linguistics Study

Every Model Has an Accent

Fırat MıhcıJune 202610 min read
TL;DR
We gave five AI models the same 102 prompts and measured only their style. The vendor is almost perfectly separable (OpenAI versus Anthropic at 96% accuracy), while the exact model lands at just 50%, with errors trapped inside each family. And the famous 'em-dash = AI' tell is really an 'em-dash = Claude' tell: Claude uses roughly five times more em-dashes than GPT.

Everyone has a theory about how to spot AI writing, and one of them is the em-dash. Too many em-dashes, the folklore goes, and a human probably did not write it. There is a real signal buried in that hunch, but it is pointed at the wrong target. This study gave five current AI models the same 102 questions and measured nothing but the style of their answers: no topic, no content words, just the texture of the prose. The result is sharp and a little surprising. You cannot reliably tell which model wrote a passage, but you can tell whose model it is with near-certainty. AI prose has an accent, and the accent belongs to the vendor, not to the individual model and not to the size of the model. The em-dash, it turns out, is one of the loudest accent markers there is, and it does not say "AI." It says "Claude."

The question: does each model have an accent?

When people say a piece of text "sounds like ChatGPT," they are claiming something testable: that a model leaves a stylistic fingerprint independent of what it is writing about. If that is true, the next question is how the fingerprints are organized. Are they per-model: does GPT-4o write differently from GPT-4o-mini? Are they per-tier: do the big flagship models share a voice that the cheaper, faster ones do not? Or are they per-vendor: does everything OpenAI ships sound like OpenAI, and everything Anthropic ships sound like Anthropic, regardless of size or generation?

These are three different worlds, and they have very different consequences for anyone trying to attribute a piece of writing. We built a study to tell them apart.

The data: same 102 prompts, five models, plain prose

The trick to measuring style rather than subject matter is to hold the subject matter still. So we built a paired corpus: five current models each answered the same 102 open-ended prompts, in plain paragraphs, at a fixed decoding temperature, in June 2026. That is 510 passages where the only thing that varies, prompt for prompt, is which model did the writing.

The five models span two vendors and three capability tiers:

  • OpenAI: GPT-4o, GPT-4o-mini, and the older GPT-3.5.
  • Anthropic: Claude Sonnet 4.5 and Claude Haiku 4.5.

We constrained every answer to ordinary prose (no bullet lists, no headings, no tables) on purpose. Formatting habits are a real and obvious tell, but they are a shallow one and easy to strip. We wanted to know whether the prose itself carries an accent once you take the layout away. So the comparison here is the hard version: paragraphs against paragraphs.

The method: 18 features, and not one of them is a content word

To measure accent without measuring topic, we described each passage with 18 content-free stylometric features: typography, rhetoric and register, sentence structure, and lexical habits. Crucially, none of them are content words. We do not look at whether a passage is about climate or cooking; we look at how long its sentences run and how much that length varies, how often it hedges, how heavily it leans on discourse connectives like however and therefore, how much passive voice it uses, how rich its vocabulary is, and (yes) how often it reaches for an em-dash or a colon.

Because the features are blind to topic, any separation they find is a separation of style. That is the whole design: if two models look different through this lens, it is because they write differently, not because they happened to be asked different things. They weren't.

The headline: the vendor is almost perfectly separable

Here is the result in one line. Asked to sort passages by vendor (OpenAI versus Anthropic), the model lands at a ROC-AUC of 0.96, which is close to categorical. From eighteen blind style features and not a single content word, you can tell an OpenAI passage from an Anthropic passage almost every time.

Now ask the harder question (which of the five models wrote a given passage), and accuracy drops to 50%. That sounds like failure until you remember the baseline: with five models, random guessing scores 20%. Fifty percent is two and a half times chance. The system clearly knows a great deal about who wrote what. It just keeps making one very specific kind of mistake.

The mistakes stay inside the family

The structure of the errors is the real finding. When the model misattributes a passage, it almost never crosses the vendor line. The two Claudes get mistaken for each other. The three GPTs get mistaken for each other. But a Claude passage is rarely mistaken for a GPT one, or the reverse.

In other words, the five models do not sit in five separate stylistic boxes. They sit in two tight vendor blocks with blurry boundaries inside each block. The vendor is a wall; the tier is a smudge. House style is something a vendor has, and every model in that vendor's lineup speaks with the house accent: the individual models are dialects of it, close enough to confuse, while the two houses stay clearly apart.

We confirmed this from the other direction too. Asked to separate tiers within a vendor, the signal nearly vanishes: telling OpenAI's three models apart from one another scores 55% against a 33% baseline, and separating Claude Sonnet from Claude Haiku lands at an AUC of just 0.65: far below the 0.96 that divides the vendors. The one model that stands out as genuinely self-identifiable is GPT-3.5, the oldest of the five, which makes sense: it is a different generation, trained by a different-era pipeline, and it kept more of its own quirks.

The em-dash points at Claude, not at "AI"

This is where the folklore meets the data. The em-dash is one of the strongest vendor discriminators in the whole feature set, but it does not separate human from machine, and it does not separate AI in general. It separates Claude from GPT. In this corpus, Claude uses roughly five times as many em-dashes as GPT does: on the order of 5.4 per thousand words versus about 1.1.

So the viral rule of thumb ("lots of em-dashes means a machine wrote it") is really tracking something much narrower. It is an "em-dash means Claude (this era)" rule wearing an "em-dash means AI" costume. Anyone who scrubs em-dashes from their drafts to look more human is, in effect, trying not to sound like one specific vendor's model from one specific year. The dash was never a property of machine writing as such. It is a property of a house style.

What each accent actually sounds like

Step back from the single dash and two distinct voices come into focus, each consistent across its vendor's models.

Claude's accent is the more punctuation-forward and rhythmically varied of the two. It uses more em-dashes and more colons, its sentence lengths swing more from short to long (higher burstiness), and it reaches for a richer, more varied vocabulary. It reads as the more performative writer, the one that varies its cadence and breaks its own sentences open with a dash.

GPT's accent is smoother and more even. Its sentences run longer and stay closer to a uniform length. It hedges more (may, might, generally) and it leans harder on discourse connectives that stitch one clause to the next. It uses more passive voice. It reads as the more cautious, more level writer, the patient explainer that keeps an even keel from sentence to sentence.

Neither of these is better writing. They are accents, in the literal sense: systematic, low-level habits that you can hear once you know to listen, and that persist no matter what the model is talking about.

Why house style is a vendor property

The cleanest way to read all of this is that prose style is set far upstream of the individual model. A vendor's training pipeline (the data it curates, the way it tunes for helpfulness, the preferences it bakes in during alignment) stamps a consistent voice onto everything that comes off the line. A flagship model and a small, fast model from the same vendor share that voice because they share that pipeline. They differ in capability; they barely differ in accent.

That is why the tier is a smudge and the vendor is a wall. Capability is a per-model property: bigger models reason better. Accent is a per-vendor property: it is a house signature, and the house is the lab that trained the family.

Why a humanizer lab studied AI accents

We build language tools. HumanizeMyAI is trained on a published corpus of 2,590 real student essays, and the skill underneath it (reading writing carefully enough to hear who is behind it) is exactly the skill this question demands. A claim like "you can tell when a machine wrote it, and here is the tell" is squarely our problem, because half the internet now edits by rules like the em-dash one. The appeal of such a rule is that it is concrete and shareable. Testing it honestly, on a paired corpus where only the author changes, with content-free features and open code, is the only way to learn whether the concrete rule is true or merely vivid. This one is half-true and aimed at the wrong target, which is worse than useless: it teaches people to launder one vendor's accent and call the result human.

This is the same instinct behind our free AI detector: read the text, not the legend. It runs through the rest of our work. Its closest sibling is It Isn't Delve, which asked whether a vocabulary list gives AI away and found the famous words to be a passing fashion; and There Is No Aesthetic Surprisal Arc, which asked whether the rhythm of surprise gives it away and found the elegant idea losing, again, to a couple of boring statistics. If you want the product-side companion on how a real-corpus approach reads these same textures, the 2,590-essay corpus deep dive is the read.

The honest caveats

This is a June-2026 snapshot of five specific model versions, and that bounds what it can claim. Accents drift as models update: the next release from either vendor can move its numbers, and a future Claude might dial its em-dashes down precisely because the tell became famous. So the durable contribution here is the reproducible method, not the fixed figures. Re-run it on next year's models and the accents will have shifted; the way to measure them will not have.

A few more limits are worth stating plainly. We fixed the decoding temperature and constrained answers to plain prose, which means we measured prose style and deliberately set formatting style aside. We collected one sample per model per prompt, not a distribution over many generations. We looked at two vendors, in English, in short-form prose. And the five-way accuracy of 50% is itself a caveat: an individual passage often is not uniquely attributable to one model. The strong, reliable signal in this study lives at the vendor level (OpenAI versus Anthropic), and the per-model and per-tier claims are correspondingly softer. We would rather under-sell that than pretend a single paragraph carries a verdict, because the entire point of the paper is that one loud feature (the em-dash) had been over-read for years.

Open data, open code

This is not a black box. The full method, every feature definition, all the figures and a frank limitations section are written up in the preprint on ResearchGate (DOI 10.13140/RG.2.2.18245.82402), and the complete analysis code together with the full 510-passage paired corpus is openly available on GitHub. We released the corpus on purpose: every model answering the same 102 prompts is a reusable instrument, and we would rather other people pull it, add their own models, and watch the accent map shift over time than ask anyone to take our numbers on faith. The em-dash was never a sign of the machine. It was a sign of the house, and we have left everything in the open so you can hear it for yourself.

About the author

Fırat Mıhcı is a computational linguist and NLP researcher. Every Model Has an Accent extends his work on AI-text detection and detector bias, grounded in a published corpus of 2,590 real student essays, the style reference inside HumanizeMyAI. Publication record at ResearchGate.

Read the full study

Complete methods, every statistic, all figures and a frank limitations section, plus the analysis code and a fully de-identified dataset.