“How many words do I need?” deserves a better answer than a random big number, and vocabulary research actually has one, or rather three, because the honest answer depends on what you want the words for. The key concept is coverage: what percentage of the running words in a text or conversation you recognise. The key threshold, from Hu and Nation’s reading research, is 98%: the coverage at which you follow comfortably and can guess the rest from context.

The three honest numbers

Everyday conversation: roughly 2,000–3,000 word families. Spoken language is dramatically more repetitive than written language. Frequency studies across languages keep finding that the top two to three thousand word families cover the large majority of everyday speech. This is why a well-chosen A2–B1 vocabulary already lets you genuinely live in Italian: order, complain, tell stories, survive relatives.

Comfortable reading of ordinary texts: somewhere around 5,000–9,000 word families. Written Italian (novels, newspapers) draws on a far longer tail of vocabulary, and reaching 98% coverage of it takes correspondingly more. This is the gap the B2–C1 years are quietly for.

Native-like breadth: tens of thousands. Educated native speakers know on the order of 15,000–20,000+ word families. You do not need this, and pretending an app will get you there in a year is marketing, not lexicography.

How Many Italian Words Do You Need? The Real Numbers

The real lesson is frequency, not the totals

The useful consequence of coverage research is about order: word value is wildly unequal. The 100 most frequent Italian words appear in essentially every paragraph you will ever read; the 7,000th word appears a few times a year. So learn by frequency, not by theme-list alphabet: the common words first, at production strength (see retrieval practice), the rare words later through volume reading, which is precisely the machine for harvesting the long tail: each rare word met in context, repeatedly, until it sticks.

Two refinements the research insists on. Count word families, not forms: parlo, parli, parlava are one family, and Italian’s rich morphology means your “known words” multiply through their forms. And knowing a word is a spectrum, from “recognise in context” to “produce with the right preposition attached”, which is why learning words inside chunks beats learning them naked.

How this shaped amo Italian

Two design decisions in amo Italian come straight from this research. First, every novel’s vocabulary is classified word by word against CEFR levels, so the books stay inside your coverage zone and the new words you meet are the right next words, frequent enough to matter. Second, the built-in dictionary holds 4,000+ curated entries with bilingual support: the conversation threshold covered with margin, deliberately, rather than an unusable full lexicon. The pre-lesson “words of the day” (see pre-learning vocabulary) then makes sure high-value words get deliberate attention before texts recycle them.

So: 2–3,000 families to talk, more to read widely, frequency as your ordering principle, and reading volume for the tail. Anyone who quotes you one big number for everything is selling something.

References and attributions: Hu & Nation, “Unknown vocabulary density and reading comprehension”, 2000; Nation, Learning Vocabulary in Another Language, 2001. Coverage figures follow Hu & Nation’s published research; paraphrased with attribution. Part of the Science of Learning Italian series.