Hebrew frequency dictionary

The Hebrew frequency dictionary ranks the 8,662 most common Hebrew words and phrases, measured across 12,933,465 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 305 Hebrew entries covers half of everything said, and the top 7,536 covers 80%. The single most common Hebrew word is “לא” (No), which accounts for 3.0% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 12.9M of speech
Distinct 332.2K word forms
Phrases 894 promoted
Glossed 8,645 to English
80.5% covered

These 8,662 entries account for 80.5% of all running Hebrew text in the corpus.

305 of them cover half of it.

Hebrew at a glance

The short answers, straight from the corpus.

Words for half of speech
305

Entries needed to cover 50% of everything said in Hebrew

Words for 80%
7,536

Entries needed to cover 80% of running Hebrew text

Most common word
לא

“No” — said 385,236 times in the Hebrew corpus

Corpus size
12,933,465 tokens

332,217 distinct Hebrew word forms were counted

Game modes

Mixed rounds

Hebrew · All three modes, shuffled

Loading Hebrew…

The ten most common Hebrew entries

Together they cover 16.8% of everything said in the corpus.

  1. 1 לא No 385.2K×
  2. 2 את You 369.4K×
  3. 3 אני I 331.4K×
  4. 4 זה This 303.4K×
  5. 5 אתה You 177.1K×
  6. 6 מה What 174.8K×
  7. 7 הוא He 118.8K×
  8. 8 לי Me 108.7K×
  9. 9 על About 104.4K×
  10. 10 כן Yes 100K×

How far each slice of the list gets you

Share of running Hebrew text covered by everything up to that rank.

305 entries reach half of all Hebrew speech.
7,536 entries reach 80% of all Hebrew speech.
Even all 8,662 entries stop short of 90% of Hebrew speech — the remainder is spread across the long tail below the cut.

The six bands, in Hebrew

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010016.33 38.4%
101–500400325.54 16.7%
501–1,000500535.12 7.0%
1,001–2,0001,0001094.80 6.5%
2,001–4,0002,0001804.47 6.1%
4,001–8,0004,0004524.15 5.8%

Word shape

Character lengths across the Hebrew top 1,000 — average 4.0.

1
2
3
4
5
6
7
8
11

Characters per word

Phrases that behave like words

894 multi-word units earned a rank of their own.

  • #79 כל כך So / So much
  • #101 אתה יודע You know
  • #113 לא יודע I don't know
  • #119 אני יודע I know
  • #142 אני חושב I think
  • #186 מה קרה What happened
  • #192 אף אחד nobody
  • #193 מה קורה what’s going on / What's happening?

What stands out in Hebrew

Read straight off this language's own numbers.

  • Just 305 entries cover half of everything said in the Hebrew subtitle corpus.
  • The ten most common Hebrew entries alone account for 16.8% of running text.
  • 894 of the top 8,662 entries are multi-word phrases that behave like single units — the most common is “כל כך”.
  • The longest single word inside the Hebrew top 1,000 is “c.hebrewrlm” (11 characters, rank 732).
  • Most of the Hebrew top 1,000 is short: 4-character words are the single largest group, and the average is 4.0 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.8% of the corpus.

How this Hebrew ranking was calculated

Built once, offline, and shipped as static data.

The source is the Hebrew side of OpenSubtitles v2024 via OPUS — 12,933,465 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Hebrew's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 332,217 distinct words appear in all; the 8,662 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 894 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 17 of the 8,662 Hebrew entries are still unresolved and are never used as game questions.

Common questions about Hebrew word frequency

Answered from this language's own corpus.

How many Hebrew words do you need to know?

Around 305 Hebrew words and phrases cover half of everything said in ordinary speech, and about 7,536 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 8,662 reaches 80.5%, so the last few thousand entries add far less than the first few hundred.

What are the most common Hebrew words?

The ten most common Hebrew entries are לא (No), את (You), אני (I), זה (This), אתה (You), מה (What), הוא (He), לי (Me), על (About) and כן (Yes). Together they account for 16.8% of all running Hebrew text in the corpus.

What is the most common word in Hebrew?

The most common Hebrew word is “לא”, meaning “No”. It appears 385,236 times across the corpus, which is 3.0% of everything said.

How is this Hebrew frequency list calculated?

The list is built from the Hebrew side of the OpenSubtitles corpus — 12,933,465 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 332,217 distinct forms found are ranked by raw count. The top 8,662 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Hebrew list include phrases as well as words?

Yes. 894 of the 8,662 Hebrew entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “כל כך” (So / So much), at rank 79. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Hebrew do the top 100 words cover?

The 100 most common Hebrew entries cover 38.4% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Hebrew frequency dictionary free to download?

Yes. All 8,645 translated Hebrew entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Hebrew's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm