Polish
frequency dictionary
The Polish frequency dictionary ranks the 9,259 most common Polish words and phrases, measured across 11,504,059 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Polish word is “nie” (no), which accounts for 3.6% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,259 entries account for 78.8% of all running Polish text in the corpus.
284 of them cover half of it.
Polish at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 284
- Most common word
- nie
- Corpus size
- 11,504,059 tokens
Entries needed to cover 50% of everything said in Polish
“no” — said 416,982 times in the Polish corpus
318,868 distinct Polish word forms were counted
Mixed rounds
Polish · All three modes, shuffled
Loading Polish…
The ten most common Polish entries
Together they cover 17.6% of everything said in the corpus.
- 1 nie no 417K×
- 2 to that 304.1K×
- 3 się myself / themselves 254.7K×
- 4 w in 189.7K×
- 5 na on 164.7K×
- 6 i and 153.9K×
- 7 co what 144.4K×
- 8 z from 138.8K×
- 9 jest is 128.5K×
- 10 że that 126.5K×
How far each slice of the list gets you
Share of running Polish text covered by everything up to that rank.
The six bands, in Polish
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 3 | 6.30 | 39.7% |
| 101–500 | 400 | 28 | 5.50 | 15.6% |
| 501–1,000 | 500 | 57 | 5.08 | 6.3% |
| 1,001–2,000 | 1,000 | 126 | 4.76 | 6.0% |
| 2,001–4,000 | 2,000 | 242 | 4.44 | 5.8% |
| 4,001–8,000 | 4,000 | 540 | 4.12 | 5.4% |
Word shape
Character lengths across the Polish top 1,000 — average 5.6.
Characters per word
Phrases that behave like words
1,177 multi-word units earned a rank of their own.
- #79 nie ma there isn't / There is none
- #87 nie wiem I don't know
- #96 w porządku okay
- #132 nie mogę I can't
- #150 nigdy nie never
- #157 po prostu just
- #182 z tobą with you
- #198 do domu home / to home
What stands out in Polish
Read straight off this language's own numbers.
- Just 284 entries cover half of everything said in the Polish subtitle corpus.
- The ten most common Polish entries alone account for 17.6% of running text.
- 1,177 of the top 9,259 entries are multi-word phrases that behave like single units — the most common is “nie ma”.
- The longest single word inside the Polish top 1,000 is “powiedziałem” (12 characters, rank 380).
- Most of the Polish top 1,000 is short: 5-character words are the single largest group, and the average is 5.6 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.4% of the corpus.
How this Polish ranking was calculated
Built once, offline, and shipped as static data.
The source is the Polish side of OpenSubtitles v2024 via OPUS — 11,504,059 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Polish's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 318,868 distinct words appear in all; the 9,259 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,177 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 92 of the 9,259 Polish entries are still unresolved and are never used as game questions.
Common questions about Polish word frequency
Answered from this language's own corpus.
What are the most common Polish words?
The ten most common Polish entries are nie (no), to (that), się (myself / themselves), w (in), na (on), i (and), co (what), z (from), jest (is) and że (that). Together they account for 17.6% of all running Polish text in the corpus.
What is the most common word in Polish?
The most common Polish word is “nie”, meaning “no”. It appears 416,982 times across the corpus, which is 3.6% of everything said.
How is this Polish frequency list calculated?
The list is built from the Polish side of the OpenSubtitles corpus — 11,504,059 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 318,868 distinct forms found are ranked by raw count. The top 9,259 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Polish list include phrases as well as words?
Yes. 1,177 of the 9,259 Polish entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “nie ma” (there isn't / There is none), at rank 79. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Polish do the top 100 words cover?
The 100 most common Polish entries cover 39.7% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Polish frequency dictionary free to download?
Yes. All 9,167 translated Polish entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Polish's.