Hungarian frequency dictionary

The Hungarian frequency dictionary ranks the 9,450 most common Hungarian words and phrases, measured across 11,865,990 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Hungarian word is “a” (the), which accounts for 6.3% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 11.9M of speech
Distinct 508.1K word forms
Phrases 1,837 promoted
Glossed 9,369 to English
76.7% covered

These 9,450 entries account for 76.7% of all running Hungarian text in the corpus.

287 of them cover half of it.

Hungarian at a glance

The short answers, straight from the corpus.

Words for half of speech
287

Entries needed to cover 50% of everything said in Hungarian

Most common word
a

“the” — said 745,972 times in the Hungarian corpus

Corpus size
11,865,990 tokens

508,133 distinct Hungarian word forms were counted

Game modes

Mixed rounds

Hungarian · All three modes, shuffled

Loading Hungarian…

The ten most common Hungarian entries

Together they cover 19.6% of everything said in the corpus.

  1. 1 a the 746K×
  2. 2 nem no 333.7K×
  3. 3 az the 283.9K×
  4. 4 hogy that 202.2K×
  5. 5 és and 174.5K×
  6. 6 egy one 162K×
  7. 7 van there is 125.4K×
  8. 8 ez this 117.6K×
  9. 9 meg start / me 92.3K×
  10. 10 de but 90.6K×

How far each slice of the list gets you

Share of running Hungarian text covered by everything up to that rank.

287 entries reach half of all Hungarian speech.
Even all 9,450 entries stop short of 80% of Hungarian speech — the remainder is spread across the long tail below the cut.
Even all 9,450 entries stop short of 90% of Hungarian speech — the remainder is spread across the long tail below the cut.

The six bands, in Hungarian

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010006.25 40.8%
101–500400375.46 14.0%
501–1,000500765.06 5.9%
1,001–2,0001,0001554.73 5.7%
2,001–4,0002,0003494.41 5.3%
4,001–8,0004,0008834.08 5.0%

Word shape

Character lengths across the Hungarian top 1,000 — average 5.4.

1
2
3
4
5
6
7
8
9
10
11
12
13

Characters per word

Phrases that behave like words

1,837 multi-word units earned a rank of their own.

  • #145 azt hiszem I think
  • #149 egy kis a little
  • #150 az egész the whole thing / the whole
  • #194 azt mondta he said
  • #197 mi történt what happened
  • #199 egy kicsit a little
  • #213 az első the first
  • #226 azt hittem I thought

What stands out in Hungarian

Read straight off this language's own numbers.

  • Just 287 entries cover half of everything said in the Hungarian subtitle corpus.
  • The ten most common Hungarian entries alone account for 19.6% of running text.
  • 1,837 of the top 9,450 entries are multi-word phrases that behave like single units — the most common is “azt hiszem”.
  • The longest single word inside the Hungarian top 1,000 is “természetesen” (13 characters, rank 572).
  • Most of the Hungarian top 1,000 is short: 5-character words are the single largest group, and the average is 5.4 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 5.0% of the corpus.

How this Hungarian ranking was calculated

Built once, offline, and shipped as static data.

The source is the Hungarian side of OpenSubtitles v2024 via OPUS — 11,865,990 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Hungarian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 508,133 distinct words appear in all; the 9,450 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,837 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 81 of the 9,450 Hungarian entries are still unresolved and are never used as game questions.

Common questions about Hungarian word frequency

Answered from this language's own corpus.

What are the most common Hungarian words?

The ten most common Hungarian entries are a (the), nem (no), az (the), hogy (that), és (and), egy (one), van (there is), ez (this), meg (start / me) and de (but). Together they account for 19.6% of all running Hungarian text in the corpus.

What is the most common word in Hungarian?

The most common Hungarian word is “a”, meaning “the”. It appears 745,972 times across the corpus, which is 6.3% of everything said.

How is this Hungarian frequency list calculated?

The list is built from the Hungarian side of the OpenSubtitles corpus — 11,865,990 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 508,133 distinct forms found are ranked by raw count. The top 9,450 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Hungarian list include phrases as well as words?

Yes. 1,837 of the 9,450 Hungarian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “azt hiszem” (I think), at rank 145. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Hungarian do the top 100 words cover?

The 100 most common Hungarian entries cover 40.8% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Hungarian frequency dictionary free to download?

Yes. All 9,369 translated Hungarian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Hungarian's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm