Italian
frequency dictionary
The Italian frequency dictionary ranks the 9,692 most common Italian words and phrases, measured across 14,553,317 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 148 Italian entries covers half of everything said, and the top 3,306 covers 80%. The single most common Italian word is “non” (not / no), which accounts for 2.5% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,692 entries account for 85.8% of all running Italian text in the corpus.
148 of them cover half of it.
Italian at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 148
- Words for 80%
- 3,306
- Most common word
- non
- Corpus size
- 14,553,317 tokens
Entries needed to cover 50% of everything said in Italian
Entries needed to cover 80% of running Italian text
“not / no” — said 369,392 times in the Italian corpus
206,152 distinct Italian word forms were counted
Mixed rounds
Italian · All three modes, shuffled
Loading Italian…
The ten most common Italian entries
Together they cover 17.9% of everything said in the corpus.
- 1 non not / no 369.4K×
- 2 di of 326.9K×
- 3 che that 325.7K×
- 4 e and 311.1K×
- 5 la the 239.3K×
- 6 il the 228.5K×
- 7 è is 227.8K×
- 8 un a / one 212.7K×
- 9 a to 201.8K×
- 10 per for 156.2K×
How far each slice of the list gets you
Share of running Italian text covered by everything up to that rank.
The six bands, in Italian
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 0 | 6.34 | 45.3% |
| 101–500 | 400 | 38 | 5.56 | 18.5% |
| 501–1,000 | 500 | 65 | 5.10 | 6.6% |
| 1,001–2,000 | 1,000 | 116 | 4.75 | 5.8% |
| 2,001–4,000 | 2,000 | 304 | 4.39 | 5.2% |
| 4,001–8,000 | 4,000 | 679 | 4.02 | 4.5% |
Word shape
Character lengths across the Italian top 1,000 — average 5.5.
Characters per word
Phrases that behave like words
1,514 multi-word units earned a rank of their own.
- #101 va bene okay / It's fine.
- #115 quello che the one who / that which
- #123 la mia mine
- #127 un po a little
- #129 lo so I know
- #135 il tuo yours
- #164 la tua yours
- #167 il suo his
What stands out in Italian
Read straight off this language's own numbers.
- Just 148 entries cover half of everything said in the Italian subtitle corpus.
- The ten most common Italian entries alone account for 17.9% of running text.
- 1,514 of the top 9,692 entries are multi-word phrases that behave like single units — the most common is “va bene”.
- The longest single word inside the Italian top 1,000 is “probabilmente” (13 characters, rank 742).
- Most of the Italian top 1,000 is short: 5-character words are the single largest group, and the average is 5.5 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.5% of the corpus.
How this Italian ranking was calculated
Built once, offline, and shipped as static data.
The source is the Italian side of OpenSubtitles v2024 via OPUS — 14,553,317 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Italian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 206,152 distinct words appear in all; the 9,692 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,514 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 268 of the 9,692 Italian entries are still unresolved and are never used as game questions.
Common questions about Italian word frequency
Answered from this language's own corpus.
How many Italian words do you need to know?
Around 148 Italian words and phrases cover half of everything said in ordinary speech, and about 3,306 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,692 reaches 85.8%, so the last few thousand entries add far less than the first few hundred.
What are the most common Italian words?
The ten most common Italian entries are non (not / no), di (of), che (that), e (and), la (the), il (the), è (is), un (a / one), a (to) and per (for). Together they account for 17.9% of all running Italian text in the corpus.
What is the most common word in Italian?
The most common Italian word is “non”, meaning “not / no”. It appears 369,392 times across the corpus, which is 2.5% of everything said.
How is this Italian frequency list calculated?
The list is built from the Italian side of the OpenSubtitles corpus — 14,553,317 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 206,152 distinct forms found are ranked by raw count. The top 9,692 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Italian list include phrases as well as words?
Yes. 1,514 of the 9,692 Italian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “va bene” (okay / It's fine.), at rank 101. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Italian do the top 100 words cover?
The 100 most common Italian entries cover 45.3% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Italian frequency dictionary free to download?
Yes. All 9,424 translated Italian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Italian's.