Russian
frequency dictionary
The Russian frequency dictionary ranks the 8,649 most common Russian words and phrases, measured across 13,108,096 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Russian word is “я” (I), which accounts for 3.0% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 8,649 entries account for 79.4% of all running Russian text in the corpus.
230 of them cover half of it.
Russian at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 230
- Most common word
- я
- Corpus size
- 13,108,096 tokens
Entries needed to cover 50% of everything said in Russian
“I” — said 398,056 times in the Russian corpus
346,185 distinct Russian word forms were counted
Mixed rounds
Russian · All three modes, shuffled
Loading Russian…
The ten most common Russian entries
Together they cover 17.7% of everything said in the corpus.
- 1 я I 398.1K×
- 2 не not / No 356.7K×
- 3 что what 283.4K×
- 4 в in 246.1K×
- 5 и and 235.9K×
- 6 ты you 229.1K×
- 7 это this 209.5K×
- 8 на on 140.4K×
- 9 с with 117.1K×
- 10 да yes 107.9K×
How far each slice of the list gets you
Share of running Russian text covered by everything up to that rank.
The six bands, in Russian
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 2 | 6.33 | 42.0% |
| 101–500 | 400 | 41 | 5.48 | 15.1% |
| 501–1,000 | 500 | 58 | 5.06 | 6.1% |
| 1,001–2,000 | 1,000 | 149 | 4.74 | 5.7% |
| 2,001–4,000 | 2,000 | 264 | 4.41 | 5.4% |
| 4,001–8,000 | 4,000 | 566 | 4.09 | 5.1% |
Word shape
Character lengths across the Russian top 1,000 — average 5.3.
Characters per word
Phrases that behave like words
1,183 multi-word units earned a rank of their own.
- #89 не знаю I don't know
- #90 у нас we / ours
- #111 я знаю I know
- #113 потому что because
- #122 не могу I can't
- #126 я хочу I want
- #130 в порядке okay
- #131 с тобой with you
What stands out in Russian
Read straight off this language's own numbers.
- Just 230 entries cover half of everything said in the Russian subtitle corpus.
- The ten most common Russian entries alone account for 17.7% of running text.
- 1,183 of the top 8,649 entries are multi-word phrases that behave like single units — the most common is “не знаю”.
- The longest single word inside the Russian top 1,000 is “действительно” (13 characters, rank 504).
- Most of the Russian top 1,000 is short: 5-character words are the single largest group, and the average is 5.3 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.1% of the corpus.
How this Russian ranking was calculated
Built once, offline, and shipped as static data.
The source is the Russian side of OpenSubtitles v2024 via OPUS — 13,108,096 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Russian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 346,185 distinct words appear in all; the 8,649 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,183 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 14 of the 8,649 Russian entries are still unresolved and are never used as game questions.
Common questions about Russian word frequency
Answered from this language's own corpus.
What are the most common Russian words?
The ten most common Russian entries are я (I), не (not / No), что (what), в (in), и (and), ты (you), это (this), на (on), с (with) and да (yes). Together they account for 17.7% of all running Russian text in the corpus.
What is the most common word in Russian?
The most common Russian word is “я”, meaning “I”. It appears 398,056 times across the corpus, which is 3.0% of everything said.
How is this Russian frequency list calculated?
The list is built from the Russian side of the OpenSubtitles corpus — 13,108,096 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 346,185 distinct forms found are ranked by raw count. The top 8,649 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Russian list include phrases as well as words?
Yes. 1,183 of the 8,649 Russian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “не знаю” (I don't know), at rank 89. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Russian do the top 100 words cover?
The 100 most common Russian entries cover 42.0% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Russian frequency dictionary free to download?
Yes. All 8,635 translated Russian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Russian's.