Which language gives you more?
Pick up to 4 languages and see them measured against each other: how many words reach half of everyday speech, how much of the language the top 100 carries, and how steeply each frequency curve falls away.
Croatian Add a language
Side by side
The best value in each row is marked — read across, not down.
| Measure | Croatian |
|---|---|
| Corpus size Tokens of subtitle dialogue counted — more data means a steadier ranking | 538.1M |
| Distinct word forms Unique surface forms found — a measure of morphology, not of quality | 1.8M |
| Words for half of speech Fewer is better — less vocabulary buys more of the language | 184 |
| Words for 80% Entries needed to cover 80% of running text | 5,640 |
| Top 100 covers Share of all speech carried by the 100 most common entries | 43.5% |
| Top 10 covers Share carried by the ten most common entries alone | 19.3% |
| Whole list covers Share of running text the full ranked list accounts for | 82.4% |
| Promoted phrases Multi-word units ranked alongside single words | 903 |
| Average word length Characters per word across the top 1,000 — descriptive, not a virtue | 5.0 |
Where each language's weight sits
Share of running text carried by each band. A front-loaded language rewards a small vocabulary more.
| Band | Croatian |
|---|---|
| Top 100 | 43.5% |
| 101–500 | 16.9% |
| 501–1,000 | 6.3% |
| 1,001–2,000 | 5.7% |
| 2,001–4,000 | 5.2% |
| 4,001–8,000 | 4.8% |
The five most common words in each
Ranked within their own corpus, translated to English.
Croatian
- 1 je is
- 2 da yes
- 3 ne no
- 4 se that
- 5 u in
Choose languages to compare
Pick up to 4.1/4 selected
The first 23 are built from the same 2,166 films, so their rankings can be compared directly.
Built from each language's own films — rankings are sound, but not directly comparable with the set above.