Persian
frequency dictionary
The Persian frequency dictionary ranks the 8,651 most common Persian words and phrases, measured across 63,944,022 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 264 Persian entries covers half of everything said, and the top 3,381 covers 80%. The single most common Persian word is “که” (that), which accounts for 2.2% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 8,651 entries account for 86.8% of all running Persian text in the corpus.
264 of them cover half of it.
Persian at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 264
- Words for 80%
- 3,381
- Most common word
- که
- Corpus size
- 63,944,022 tokens
Entries needed to cover 50% of everything said in Persian
Entries needed to cover 80% of running Persian text
“that” — said 1,396,273 times in the Persian corpus
778,409 distinct Persian word forms were counted
Mixed rounds
Persian · All three modes, shuffled
Loading Persian…
The ten most common Persian entries
Together they cover 15.7% of everything said in the corpus.
- 1 که that 1.4M×
- 2 رو on / Ro 1.3M×
- 3 من I 1.2M×
- 4 به to 1.1M×
- 5 و and 1.1M×
- 6 تو you 940K×
- 7 از from 936.4K×
- 8 اون that 709.5K×
- 9 اين this 663.3K×
- 10 يه a / One 645.7K×
How far each slice of the list gets you
Share of running Persian text covered by everything up to that rank.
The six bands, in Persian
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 0 | 6.31 | 37.8% |
| 101–500 | 400 | 12 | 5.63 | 20.5% |
| 501–1,000 | 500 | 30 | 5.22 | 8.6% |
| 1,001–2,000 | 1,000 | 65 | 4.88 | 7.9% |
| 2,001–4,000 | 2,000 | 178 | 4.50 | 6.7% |
| 4,001–8,000 | 4,000 | 381 | 4.09 | 5.3% |
Word shape
Character lengths across the Persian top 1,000 — average 4.1.
Characters per word
Phrases that behave like words
725 multi-word units earned a rank of their own.
- #198 در مورد About
- #240 بعد از After
- #250 به نظر It seems
- #254 به خاطر because / Because of
- #256 اينه که that's why / That's it.
- #259 قبل از Before
- #314 صبر کن Wait
- #336 خداي من Oh my God
What stands out in Persian
Read straight off this language's own numbers.
- Just 264 entries cover half of everything said in the Persian subtitle corpus.
- The ten most common Persian entries alone account for 15.7% of running text.
- 725 of the top 8,651 entries are multi-word phrases that behave like single units — the most common is “در مورد”.
- The longest single word inside the Persian top 1,000 is “اميدوارم” (8 characters, rank 679).
- Most of the Persian top 1,000 is short: 4-character words are the single largest group, and the average is 4.1 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.3% of the corpus.
How this Persian ranking was calculated
Built once, offline, and shipped as static data.
The source is the Persian side of OpenSubtitles v2018 via OPUS — 63,944,022 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Persian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 778,409 distinct words appear in all; the 8,651 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 725 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 14 of the 8,651 Persian entries are still unresolved and are never used as game questions.
Common questions about Persian word frequency
Answered from this language's own corpus.
How many Persian words do you need to know?
Around 264 Persian words and phrases cover half of everything said in ordinary speech, and about 3,381 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 8,651 reaches 86.8%, so the last few thousand entries add far less than the first few hundred.
What are the most common Persian words?
The ten most common Persian entries are که (that), رو (on / Ro), من (I), به (to), و (and), تو (you), از (from), اون (that), اين (this) and يه (a / One). Together they account for 15.7% of all running Persian text in the corpus.
What is the most common word in Persian?
The most common Persian word is “که”, meaning “that”. It appears 1,396,273 times across the corpus, which is 2.2% of everything said.
How is this Persian frequency list calculated?
The list is built from the Persian side of the OpenSubtitles corpus — 63,944,022 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 778,409 distinct forms found are ranked by raw count. The top 8,651 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Persian list include phrases as well as words?
Yes. 725 of the 8,651 Persian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “در مورد” (About), at rank 198. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Persian do the top 100 words cover?
The 100 most common Persian entries cover 37.8% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Persian frequency dictionary free to download?
Yes. All 8,637 translated Persian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Persian's.