Turkish
frequency dictionary
The Turkish frequency dictionary ranks the 9,375 most common Turkish words and phrases, measured across 10,892,222 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Turkish word is “bir” (one), which accounts for 2.5% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,375 entries account for 75.1% of all running Turkish text in the corpus.
606 of them cover half of it.
Turkish at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 606
- Most common word
- bir
- Corpus size
- 10,892,222 tokens
Entries needed to cover 50% of everything said in Turkish
“one” — said 273,814 times in the Turkish corpus
456,733 distinct Turkish word forms were counted
Mixed rounds
Turkish · All three modes, shuffled
Loading Turkish…
The ten most common Turkish entries
Together they cover 10.7% of everything said in the corpus.
- 1 bir one 273.8K×
- 2 bu this 171.5K×
- 3 ne what 128.6K×
- 4 ve and 119.4K×
- 5 için for 80.4K×
- 6 mi is it / me 79.3K×
- 7 çok very 79K×
- 8 o he 78.8K×
- 9 evet yes 76.9K×
- 10 ben I 76.6K×
How far each slice of the list gets you
Share of running Turkish text covered by everything up to that rank.
The six bands, in Turkish
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 1 | 6.29 | 31.1% |
| 101–500 | 400 | 25 | 5.55 | 16.9% |
| 501–1,000 | 500 | 37 | 5.13 | 7.1% |
| 1,001–2,000 | 1,000 | 66 | 4.82 | 6.9% |
| 2,001–4,000 | 2,000 | 143 | 4.51 | 6.7% |
| 4,001–8,000 | 4,000 | 300 | 4.18 | 6.4% |
Word shape
Character lengths across the Turkish top 1,000 — average 5.5.
Characters per word
Phrases that behave like words
657 multi-word units earned a rank of their own.
- #38 bir şey something
- #195 teşekkür ederim thank you
- #213 biliyor musun do you know
- #225 özür dilerim I’m sorry / I apologize
- #231 her şeyi everything
- #236 hiçbir şey nothing
- #239 bu yüzden that’s why / Therefore
- #240 ne yapıyorsun what are you doing
What stands out in Turkish
Read straight off this language's own numbers.
- Just 606 entries cover half of everything said in the Turkish subtitle corpus.
- The ten most common Turkish entries alone account for 10.7% of running text.
- 657 of the top 9,375 entries are multi-word phrases that behave like single units — the most common is “bir şey”.
- The longest single word inside the Turkish top 1,000 is “hissediyorum” (12 characters, rank 648).
- Most of the Turkish top 1,000 is short: 5-character words are the single largest group, and the average is 5.5 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 6.4% of the corpus.
How this Turkish ranking was calculated
Built once, offline, and shipped as static data.
The source is the Turkish side of OpenSubtitles v2024 via OPUS — 10,892,222 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Turkish's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 456,733 distinct words appear in all; the 9,375 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 657 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 253 of the 9,375 Turkish entries are still unresolved and are never used as game questions.
Common questions about Turkish word frequency
Answered from this language's own corpus.
What are the most common Turkish words?
The ten most common Turkish entries are bir (one), bu (this), ne (what), ve (and), için (for), mi (is it / me), çok (very), o (he), evet (yes) and ben (I). Together they account for 10.7% of all running Turkish text in the corpus.
What is the most common word in Turkish?
The most common Turkish word is “bir”, meaning “one”. It appears 273,814 times across the corpus, which is 2.5% of everything said.
How is this Turkish frequency list calculated?
The list is built from the Turkish side of the OpenSubtitles corpus — 10,892,222 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 456,733 distinct forms found are ranked by raw count. The top 9,375 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Turkish list include phrases as well as words?
Yes. 657 of the 9,375 Turkish entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “bir şey” (something), at rank 38. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Turkish do the top 100 words cover?
The 100 most common Turkish entries cover 31.1% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Turkish frequency dictionary free to download?
Yes. All 9,122 translated Turkish entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Turkish's.