Croatian frequency dictionary

The Croatian frequency dictionary ranks the 8,857 most common Croatian words and phrases, measured across 538,148,022 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 184 Croatian entries covers half of everything said, and the top 5,640 covers 80%. The single most common Croatian word is “je” (is), which accounts for 3.9% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 538.1M of speech
Distinct 1.8M word forms
Phrases 903 promoted
Glossed 8,686 to English
82.4% covered

These 8,857 entries account for 82.4% of all running Croatian text in the corpus.

184 of them cover half of it.

Croatian at a glance

The short answers, straight from the corpus.

Words for half of speech
184

Entries needed to cover 50% of everything said in Croatian

Words for 80%
5,640

Entries needed to cover 80% of running Croatian text

Most common word
je

“is” — said 21,028,100 times in the Croatian corpus

Corpus size
538,148,022 tokens

1,764,942 distinct Croatian word forms were counted

Game modes

Mixed rounds

Croatian · All three modes, shuffled

Loading Croatian…

The ten most common Croatian entries

Together they cover 19.3% of everything said in the corpus.

  1. 1 je is 21M×
  2. 2 da yes 16.6M×
  3. 3 ne no 10.5M×
  4. 4 se that 10.2M×
  5. 5 u in 9.1M×
  6. 6 i and 9.1M×
  7. 7 to that 8M×
  8. 8 sam am / I am 7.4M×
  9. 9 što what 6.4M×
  10. 10 na on 5.3M×

How far each slice of the list gets you

Share of running Croatian text covered by everything up to that rank.

184 entries reach half of all Croatian speech.
5,640 entries reach 80% of all Croatian speech.
Even all 8,857 entries stop short of 90% of Croatian speech — the remainder is spread across the long tail below the cut.

The six bands, in Croatian

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010036.34 43.5%
101–500400305.55 16.9%
501–1,000500335.07 6.3%
1,001–2,0001,000954.73 5.7%
2,001–4,0002,0001984.39 5.2%
4,001–8,0004,0004354.06 4.8%

Word shape

Character lengths across the Croatian top 1,000 — average 5.0.

1
2
3
4
5
6
7
8
9
10
11
12

Characters per word

Phrases that behave like words

903 multi-word units earned a rank of their own.

  • #50 u redu okay
  • #82 mislim da I think that / I think
  • #100 ne znam I don't know
  • #101 ne mogu I can't
  • #107 zar ne right? / isn't it?
  • #119 jesi li are you
  • #171 žao mi je I'm sorry
  • #186 ono što what

What stands out in Croatian

Read straight off this language's own numbers.

  • Just 184 entries cover half of everything said in the Croatian subtitle corpus.
  • The ten most common Croatian entries alone account for 19.3% of running text.
  • 903 of the top 8,857 entries are multi-word phrases that behave like single units — the most common is “u redu”.
  • The longest single word inside the Croatian top 1,000 is “nevjerojatno” (12 characters, rank 903).
  • Most of the Croatian top 1,000 is short: 5-character words are the single largest group, and the average is 5.0 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.8% of the corpus.

How this Croatian ranking was calculated

Built once, offline, and shipped as static data.

The source is the Croatian side of OpenSubtitles v2018 via OPUS — 538,148,022 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Croatian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 1,764,942 distinct words appear in all; the 8,857 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 903 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 171 of the 8,857 Croatian entries are still unresolved and are never used as game questions.

Common questions about Croatian word frequency

Answered from this language's own corpus.

How many Croatian words do you need to know?

Around 184 Croatian words and phrases cover half of everything said in ordinary speech, and about 5,640 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 8,857 reaches 82.4%, so the last few thousand entries add far less than the first few hundred.

What are the most common Croatian words?

The ten most common Croatian entries are je (is), da (yes), ne (no), se (that), u (in), i (and), to (that), sam (am / I am), što (what) and na (on). Together they account for 19.3% of all running Croatian text in the corpus.

What is the most common word in Croatian?

The most common Croatian word is “je”, meaning “is”. It appears 21,028,100 times across the corpus, which is 3.9% of everything said.

How is this Croatian frequency list calculated?

The list is built from the Croatian side of the OpenSubtitles corpus — 538,148,022 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 1,764,942 distinct forms found are ranked by raw count. The top 8,857 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Croatian list include phrases as well as words?

Yes. 903 of the 8,857 Croatian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “u redu” (okay), at rank 50. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Croatian do the top 100 words cover?

The 100 most common Croatian entries cover 43.5% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Croatian frequency dictionary free to download?

Yes. All 8,686 translated Croatian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Croatian's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm