Czech frequency dictionary

The Czech frequency dictionary ranks the 9,145 most common Czech words and phrases, measured across 12,454,905 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 233 Czech entries covers half of everything said, and the top 6,164 covers 80%. The single most common Czech word is “to” (that), which accounts for 3.5% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 12.5M of speech
Distinct 329.6K word forms
Phrases 931 promoted
Glossed 9,056 to English
81.8% covered

These 9,145 entries account for 81.8% of all running Czech text in the corpus.

233 of them cover half of it.

Czech at a glance

The short answers, straight from the corpus.

Words for half of speech
233

Entries needed to cover 50% of everything said in Czech

Words for 80%
6,164

Entries needed to cover 80% of running Czech text

Most common word
to

“that” — said 440,493 times in the Czech corpus

Corpus size
12,454,905 tokens

329,646 distinct Czech word forms were counted

Game modes

Mixed rounds

Czech · All three modes, shuffled

Loading Czech…

The ten most common Czech entries

Together they cover 17.2% of everything said in the corpus.

  1. 1 to that 440.5K×
  2. 2 se with 292.1K×
  3. 3 je is 288.2K×
  4. 4 a and 227.3K×
  5. 5 na on 180.4K×
  6. 6 jsem I am 170.8K×
  7. 7 že that 158.1K×
  8. 8 co what 157K×
  9. 9 v in 121.3K×
  10. 10 si yourself 111.4K×

How far each slice of the list gets you

Share of running Czech text covered by everything up to that rank.

233 entries reach half of all Czech speech.
6,164 entries reach 80% of all Czech speech.
Even all 9,145 entries stop short of 90% of Czech speech — the remainder is spread across the long tail below the cut.

The six bands, in Czech

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010006.38 40.9%
101–500400225.58 17.5%
501–1,000500365.11 6.8%
1,001–2,0001,0001094.76 6.1%
2,001–4,0002,0002104.43 5.6%
4,001–8,0004,0004364.08 5.1%

Word shape

Character lengths across the Czech top 1,000 — average 4.9.

1
2
3
4
5
6
7
8
9
10
11
12

Characters per word

Phrases that behave like words

931 multi-word units earned a rank of their own.

  • #106 v pořádku okay
  • #113 no tak come on
  • #129 myslím že I think that
  • #162 o tom about that
  • #199 s tím with that
  • #207 se mnou with me
  • #236 s tebou with you
  • #297 v pohodě okay / It's fine.

What stands out in Czech

Read straight off this language's own numbers.

  • Just 233 entries cover half of everything said in the Czech subtitle corpus.
  • The ten most common Czech entries alone account for 17.2% of running text.
  • 931 of the top 9,145 entries are multi-word phrases that behave like single units — the most common is “v pořádku”.
  • The longest single word inside the Czech top 1,000 is “poslouchejte” (12 characters, rank 705).
  • Most of the Czech top 1,000 is short: 5-character words are the single largest group, and the average is 4.9 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.1% of the corpus.

How this Czech ranking was calculated

Built once, offline, and shipped as static data.

The source is the Czech side of OpenSubtitles v2024 via OPUS — 12,454,905 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Czech's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 329,646 distinct words appear in all; the 9,145 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 931 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 89 of the 9,145 Czech entries are still unresolved and are never used as game questions.

Common questions about Czech word frequency

Answered from this language's own corpus.

How many Czech words do you need to know?

Around 233 Czech words and phrases cover half of everything said in ordinary speech, and about 6,164 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,145 reaches 81.8%, so the last few thousand entries add far less than the first few hundred.

What are the most common Czech words?

The ten most common Czech entries are to (that), se (with), je (is), a (and), na (on), jsem (I am), že (that), co (what), v (in) and si (yourself). Together they account for 17.2% of all running Czech text in the corpus.

What is the most common word in Czech?

The most common Czech word is “to”, meaning “that”. It appears 440,493 times across the corpus, which is 3.5% of everything said.

How is this Czech frequency list calculated?

The list is built from the Czech side of the OpenSubtitles corpus — 12,454,905 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 329,646 distinct forms found are ranked by raw count. The top 9,145 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Czech list include phrases as well as words?

Yes. 931 of the 9,145 Czech entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “v pořádku” (okay), at rank 106. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Czech do the top 100 words cover?

The 100 most common Czech entries cover 40.9% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Czech frequency dictionary free to download?

Yes. All 9,056 translated Czech entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Czech's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm