Arabic frequency dictionary

The Arabic frequency dictionary ranks the 8,685 most common Arabic words and phrases, measured across 11,806,813 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Arabic word is “لا” (No), which accounts for 2.0% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 11.8M of speech
Distinct 549K word forms
Phrases 735 promoted
Glossed 8,638 to English
72.9% covered

These 8,685 entries account for 72.9% of all running Arabic text in the corpus.

745 of them cover half of it.

Arabic at a glance

The short answers, straight from the corpus.

Words for half of speech
745

Entries needed to cover 50% of everything said in Arabic

Most common word
لا

“No” — said 234,495 times in the Arabic corpus

Corpus size
11,806,813 tokens

549,006 distinct Arabic word forms were counted

Game modes

Mixed rounds

Arabic · All three modes, shuffled

Loading Arabic…

The ten most common Arabic entries

Together they cover 12.4% of everything said in the corpus.

  1. 1 لا No 234.5K×
  2. 2 من From 202.5K×
  3. 3 في In 174.4K×
  4. 4 أن That 164K×
  5. 5 هذا This 162.7K×
  6. 6 يا oh 118.2K×
  7. 7 ما What 109.9K×
  8. 8 على On / Ali 106.8K×
  9. 9 هل Is / Do you 99.6K×
  10. 10 أنا I 88.7K×

How far each slice of the list gets you

Share of running Arabic text covered by everything up to that rank.

745 entries reach half of all Arabic speech.
Even all 8,685 entries stop short of 80% of Arabic speech — the remainder is spread across the long tail below the cut.
Even all 8,685 entries stop short of 90% of Arabic speech — the remainder is spread across the long tail below the cut.

The six bands, in Arabic

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010016.23 31.7%
101–500400335.50 14.6%
501–1,000500375.10 6.6%
1,001–2,0001,000924.81 6.8%
2,001–4,0002,0001674.51 6.8%
4,001–8,0004,0003374.20 6.6%

Word shape

Character lengths across the Arabic top 1,000 — average 4.1.

1
2
3
4
5
6
7
8

Characters per word

Phrases that behave like words

735 multi-word units earned a rank of their own.

  • #48 يجب أن Must / It must
  • #135 أليس كذلك isn’t that right / Isn't that right?
  • #138 يا إلهي Oh my God / Oh God
  • #162 أريد أن I want to
  • #165 لم يكن wasn’t / It was not
  • #174 يمكن أن could / It can
  • #177 من قبل before / By
  • #178 يا رجل man / Oh man

What stands out in Arabic

Read straight off this language's own numbers.

  • Just 745 entries cover half of everything said in the Arabic subtitle corpus.
  • The ten most common Arabic entries alone account for 12.4% of running text.
  • 735 of the top 8,685 entries are multi-word phrases that behave like single units — the most common is “يجب أن”.
  • The longest single word inside the Arabic top 1,000 is “بالتأكيد” (8 characters, rank 188).
  • Most of the Arabic top 1,000 is short: 4-character words are the single largest group, and the average is 4.1 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 6.6% of the corpus.

How this Arabic ranking was calculated

Built once, offline, and shipped as static data.

The source is the Arabic side of OpenSubtitles v2024 via OPUS — 11,806,813 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Arabic's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 549,006 distinct words appear in all; the 8,685 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 735 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 47 of the 8,685 Arabic entries are still unresolved and are never used as game questions.

Common questions about Arabic word frequency

Answered from this language's own corpus.

What are the most common Arabic words?

The ten most common Arabic entries are لا (No), من (From), في (In), أن (That), هذا (This), يا (oh), ما (What), على (On / Ali), هل (Is / Do you) and أنا (I). Together they account for 12.4% of all running Arabic text in the corpus.

What is the most common word in Arabic?

The most common Arabic word is “لا”, meaning “No”. It appears 234,495 times across the corpus, which is 2.0% of everything said.

How is this Arabic frequency list calculated?

The list is built from the Arabic side of the OpenSubtitles corpus — 11,806,813 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 549,006 distinct forms found are ranked by raw count. The top 8,685 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Arabic list include phrases as well as words?

Yes. 735 of the 8,685 Arabic entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “يجب أن” (Must / It must), at rank 48. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Arabic do the top 100 words cover?

The 100 most common Arabic entries cover 31.7% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Arabic frequency dictionary free to download?

Yes. All 8,638 translated Arabic entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Arabic's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm