Ask how someone's English gets better and you will hear a familiar story: they learn longer words, build longer sentences, fold in more clauses. It is an intuitive picture, and it is mostly wrong. We took 3,300 learner essays graded from beginner to near-native, plus a small set of essays by native writers, and measured what actually changes as English proficiency grows. One axis dominates everything: accuracy. Errors fall almost five-fold across the ladder and, at the top, land right on the native baseline. Two quieter changes ride alongside it: vocabulary slowly gets rarer, and writing gets structurally more elaborate. But the elaboration is phrasal, not clausal: advanced writers pack more into their noun phrases, they do not pile up subordinate clauses. And the one error that refuses to go away is the humblest of all: punctuation.
The data: 3,300 essays, six levels, one public corpus
To ask how writing changes with proficiency you need a lot of real writing, graded consistently, by writers across the whole range. The W&I+LOCNESS corpus (the public dataset behind the 2019 BEA shared task on grammatical error correction) is built for exactly that. It collects 3,300 learner essays labelled by CEFR level, the six-band scale that runs A1 (beginner) → A2 → B1 → B2 → C1 → C2 (near-native), plus 50 essays by native English writers as a reference point. Crucially, each essay carries gold human error annotations as well as the text itself: a trained annotator has marked every error and its type. That lets us measure accuracy directly, from human judgements, rather than guessing at it with a machine.
One caution sits over everything that follows, so we put it up front. This is cross-sectional, not longitudinal: the essays at each level are written by different people. We are watching how these measures co-vary with proficiency across many writers, not following one writer as they improve over years. That is a real limit, and we say so plainly throughout.
Finding 1: accuracy is the dominant axis, and it converges on natives
The strongest, cleanest trend of every measure we computed is error density, and it points one way: down, steeply. Beginner (A1) essays carry about 20.7 errors per 100 words. Near-native (C2) essays carry about 4.0. That is an almost five-fold drop, and it is the single most pronounced pattern in the data.
What makes it more than a trend is where it lands. The native essays sit at roughly 5.0 errors per 100 words, and by C2 the learners have essentially converged on that baseline. (Yes, the most advanced learners score slightly cleaner than the natives, but the native set is small and a different genre, so we read that as "they have caught up," not "they have overtaken.") If you had to name what "getting better at English" means in this corpus, in one number, this is it: you stop making errors, until you make about as few as a native writer does.
Finding 2: errors don't fade evenly, and punctuation gets left behind
Accuracy improves overall, but not uniformly across error types, and the uneven part is the interesting part. We grouped the CEFR levels into broad beginner (A), intermediate (B) and advanced (C) bands and watched each error category shrink at its own pace.
| Error type | How much it shrinks (beginner → advanced) |
|---|---|
| Orthography / capitalization | ~8× (collapses fastest) |
| Spelling | ~4× |
| Verb tense | ~4× |
| Prepositions | ~2.5× (stubborn) |
| Punctuation | ~2× (most stubborn) |
The pattern is clean. The surface-mechanical mistakes (getting letters and capitals right, spelling words correctly, choosing the right tense) fall away fastest, because they are the most teachable and the most rule-bound. What lingers are the judgement-heavy choices: which preposition a verb takes, where a comma belongs. These do not reduce to a rule you memorize; they take a feel for the language that arrives late.
There is a striking consequence. Because the total error budget shrinks so much while punctuation shrinks so little, punctuation becomes the single largest error category for advanced learners. It rises from about 15% of a beginner's errors to roughly 22% of an advanced learner's, not because advanced writers punctuate worse, but because they have fixed almost everything else. For a teacher, that is an oddly precise piece of guidance: with a strong student, the highest-yield place to look is the comma.
Finding 3: complexity grows in the noun phrase, not the clause
Here is where the intuitive story breaks. Writing does get more structurally complex as proficiency rises, but only along one of two possible routes. We tracked noun-phrase modification (how much detail gets packed around nouns: adjectives, prepositional phrases, relative clauses) and average dependency distance (a measure of how far apart grammatically-linked words sit, which rises as sentences get more intricately wired). Both climb steadily with proficiency. But clausal subordination (stacking subordinate clauses inside one another, the thing most people picture when they say "complex sentence") is essentially flat. It is the weakest trend of everything we measured.
This is not a fluke of one dataset. It is an independent confirmation of a well-known argument by Biber, Gray and Poonpon (2011), who proposed that advanced academic writing grows dense not by subordination but by phrasal compression: loading information into elaborated noun phrases. Their claim was based on register comparison; our data shows the same shape emerging across learner development. Mature writing says more per noun, not more clauses per sentence.
One popular shortcut takes friendly fire here. Mean sentence length (the go-to proxy for "complexity" in countless rubrics and tools) barely tracks learner proficiency at all in this data. It only jumps once you reach the native essays, and stays nearly flat across the entire A1-to-C2 learner range. A writer can grow enormously in skill without their sentences getting meaningfully longer. If you are judging development by sentence length, you are mostly measuring nothing.
A steady second axis: words get rarer
Alongside accuracy, vocabulary does its own slow work. Diversity rises (advanced writers reach for a wider range of words), and the words themselves get rarer: the share of lower-frequency words climbs from about 10% to 16% across the ladder. It is a real, monotonic signal, just a gentler one than accuracy. Proficiency reads partly as precision (fewer errors) and partly as range (less reliance on the most common few thousand words).
Why a humanizer lab studied learner English
We build language tools. HumanizeMyAI is trained on a published corpus of 2,590 real student essays, and the skill underneath it (reading human writing carefully, across levels and registers, and noticing what actually distinguishes one kind from another) is exactly the skill this study asks for. A claim like "longer sentences mean better writing" is squarely our problem, because so much language teaching, and so much automated scoring, quietly runs on it.
There is a sharper reason too, and it is uncomfortable. AI detectors disproportionately misflag non-native English writers. The same careful, formal, slightly-conservative prose that earns a high CEFR band (even, controlled, low in surprise) is the texture detectors mistake for machine output, and second-language students pay for it. Understanding what advanced learner writing really looks like is part of why we built our free AI detector to be read with humility rather than treated as a verdict. The same instinct drove the two siblings to this study: It Isn't Delve, which found that the famous "AI vocabulary" is a passing fashion rather than a reliable tell, and There Is No Aesthetic Surprisal Arc, which found that the rhythm of writing doesn't give AI away: only its flatness does, the very flatness a fluent non-native writer can have honestly.
The honest caveats (which shape how to read this)
We would rather under-claim, so here is exactly what this study does and does not show.
It is cross-sectional. Different writers sit at each level, so we observe how proficiency co-varies with these measures across people, not one person changing over time. A genuine longitudinal study could, in principle, find a different developmental path inside individuals.
The CEFR labels are holistic (a single overall judgement per essay, not a precise measurement), so the bands are coarse by nature. The native baseline is small (just 50 essays) and a somewhat different genre, which is why we treat the native comparisons as suggestive rather than settled. And it is a single corpus: a clean, well-annotated, public one, but one source. The honest reading of every number here is "this is what the data shows, robustly, in this corpus," not "this is a law of English." The point of putting the code and data in the open is that anyone can test how far it travels.
Open data, open code
This is not a black box. The full method, every measure, all the figures and a frank limitations section are written up in the preprint on ResearchGate (DOI 10.13140/RG.2.2.35861.90086), and the complete analysis code with a one-command reproduction is openly available on GitHub. The dataset itself (W&I+LOCNESS) is public, so anyone can pull the same 3,300 essays, run the same measures, and check the result rather than take it on faith. The tidy story that better English means longer, more clause-heavy sentences does not survive contact with the data; what survives is duller and more useful: accuracy converging on natives, complexity growing in the noun phrase, and the comma quietly outlasting every other mistake.
About the author
Fırat Mıhcı is a computational linguist and NLP researcher. What Actually Changes as English Proficiency Grows extends his work on AI-text detection and detector bias, grounded in a published corpus of 2,590 real student essays, the style reference inside HumanizeMyAI. Publication record at ResearchGate.
Read the full study
Complete methods, every statistic, all figures and a frank limitations section, plus the analysis code and a fully de-identified dataset.