How many words do you need to read a novel? We measured 14 classics
Original data, computed 2026-07-26. Method and caveats below — short version: we counted, we didn’t guess.
Ask this question in a language forum and you’ll get the folklore numbers: “3,000 words is conversational”, “8,000 to read novels”. So we measured. We ran fourteen public-domain classics in six languages through the same frequency-and-lemma pipeline that powers the Glossa reader and asked, for each book: if you knew the N most frequent word families of the language, how much of this book’s running text would you recognise?
The numbers
Coverage = share of the running text you’d recognise. The last two columns show the vocabulary size where coverage crosses 95% (followable with effort) and 98% (comfortable) — see the coverage math for why those two lines matter. “—” means the threshold isn’t reached within the 50,000 most frequent families of the modern language.
| Book | Language | Know 2,000 families | Know 5,000 | Families for 95% | For 98% |
|---|---|---|---|---|---|
| Buddenbrooks — Thomas Mann, 1901 | German | 76% | 83% | 33,957 | — |
| Die Verwandlung — Franz Kafka, 1915 | German | 82% | 89% | 17,888 | — |
| Pride and Prejudice — Jane Austen, 1813 | English | 85% | 92% | 10,842 | 25,741 |
| Don Quijote — Miguel de Cervantes, 1605 | Spanish | 82% | 88% | 25,970 | — |
| Insolación y Morriña — Emilia Pardo Bazán, 1889 | Spanish | 75% | 84% | — | — |
| Niebla — Miguel de Unamuno, 1914 | Spanish | 84% | 90% | 12,899 | — |
| Les trois mousquetaires — Alexandre Dumas, 1844 | French | 84% | 90% | 17,662 | 31,450 |
| Les Misérables (T. I: Fantine) — Victor Hugo, 1862 | French | 83% | 90% | 12,879 | — |
| Voyage au centre de la Terre — Jules Verne, 1864 | French | 81% | 89% | 12,205 | 30,402 |
| Du côté de chez Swann — Marcel Proust, 1913 | French | 85% | 91% | 11,189 | 28,311 |
| Le avventure di Pinocchio — Carlo Collodi, 1883 | Italian | 78% | 85% | — | — |
| L'Innocente — Gabriele D'Annunzio, 1892 | Italian | 76% | 85% | 29,991 | — |
| Os Maias — Eça de Queirós, 1888 * | Portuguese | 69% | 76% | — | — |
| Dom Casmurro — Machado de Assis, 1899 * | Portuguese | 78% | 84% | — | — |
* Source edition uses pre-reform Portuguese spelling, which inflates these figures — see the finding on editions below. Proper nouns are excluded throughout (knowing “Quijote” is not vocabulary).
What the data says
1. The folklore numbers are optimistic. Two thousand word families — a solid course-app vocabulary — buys you 69–85% coverage of a real classic: one unknown word in every three to seven. That’s decrypting, not reading. Even 5,000 families tops out at 76–92%. The comfortable-reading line (98%) sits at 25,000+ families where it’s reachable at all — a vocabulary nobody builds from study alone. The honest conclusion: for classic literature, you don’t read novels because your vocabulary is big enough; your vocabulary gets big enough because you read novels — starting before the numbers say you’re ready, with support.
2. The book matters more than the language. Within Spanish, Unamuno’s Niebla reaches 95% at ~12,900 families while Don Quijote needs ~26,000 — twice the vocabulary, same language, exactly the Golden-Age-Spanish warning from our Spanish guide. In German, Kafka’s Die Verwandlung (~17,900) undercuts Mann’s Buddenbrooks (~34,000) by nearly half — “Thomas Mann is a wall at B2”, now with a number on it.
3. Proust’s vocabulary is not the problem. The counterintuitive winner: Du côté de chez Swann needs fewer word families for 95% (~11,200) than The Three Musketeers (~17,700). Proust’s famous difficulty lives in his sentences, not his lexicon — the same lesson as Dostoevsky in our Russian guide: word counts measure one axis of difficulty, and syntax is another.
4. The edition can cost more than the author. Our two Portuguese classics look like the hardest books in the table — but much of that is spelling: the free editions use pre-reform Portuguese orthography, so thousands of tokens (elle, aquella, pharmacia) miss the modern frequency list entirely. It’s the same trap as German Fraktur and Russia’s pre-1918 letters, quantified: check the edition before you blame yourself.
5. The graded-reader cliff is real and measurable. Top graded-reader series stop at ~3,000 headwords. Knowing 3,000 families covers 72–88% of the classics here — squarely in the decrypting zone — the exact gap the graded-readers article describes, now in columns.
Method (and what to make of the absolute numbers)
Full texts from Project Gutenberg, boilerplate stripped. Words are lemmatised (so was/were/been all count as be) and matched against the language’s word families ranked by aggregated corpus frequency (wordfreq). Proper nouns are excluded; hyphenated compounds count by their parts; German compounds are credited when both parts are known — mirroring how our reader’s difficulty pipeline treats them. Where “—” appears, more than ~5% of a book’s tokens fall outside the modern 50,000-family list — archaic spelling, dialect, or regional diminutives — so the threshold stops being meaningful. And one honest caveat: “word family” has no universal definition, so compare rows against each other, not against figures from studies that count families differently. The relative story — which books, which editions, which authors — is the robust part.
So what do you do with a 5,000-word vocabulary?
Read anyway — with the gap covered. The whole point of adaptive interlinear reading is that the 10–20% of words you don’t know arrive with their meaning already underneath, so a 84%-coverage book reads like a 98% one while the middle of the table does its work. That’s what Glossa does with any book you upload and the classics in its catalog — including most of the books in this table. The free tier (two books, 10 pages) is enough to feel what 98% effective coverage is like on a book “above your level”.
