Korean frequency dictionary

The Korean frequency dictionary ranks the 9,026 most common Korean words and phrases, measured across 8,360,774 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Korean word is “내가” (I), which accounts for 0.5% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 8.4M of speech
Distinct 722.8K word forms
Phrases 381 promoted
Glossed 8,835 to English
65.5% covered

These 9,026 entries account for 65.5% of all running Korean text in the corpus.

1,738 of them cover half of it.

Korean at a glance

The short answers, straight from the corpus.

Words for half of speech
1,738

Entries needed to cover 50% of everything said in Korean

Most common word
내가

“I” — said 45,864 times in the Korean corpus

Corpus size
8,360,774 tokens

722,759 distinct Korean word forms were counted

Game modes

Mixed rounds

Korean · All three modes, shuffled

Loading Korean…

The ten most common Korean entries

Together they cover 4.5% of everything said in the corpus.

  1. 1 내가 I 45.9K×
  2. 2 I 41.1K×
  3. 3 That 40.9K×
  4. 4 No 39.7K×
  5. 5 My / Mine 39.2K×
  6. 6 Number / Thursday 38.6K×
  7. 7 This 35.8K×
  8. 8 Yes 34.2K×
  9. 9 거야 That / That's it. 33.5K×
  10. 10 you you 29.9K×

How far each slice of the list gets you

Share of running Korean text covered by everything up to that rank.

1,738 entries reach half of all Korean speech.
Even all 9,026 entries stop short of 80% of Korean speech — the remainder is spread across the long tail below the cut.
Even all 9,026 entries stop short of 90% of Korean speech — the remainder is spread across the long tail below the cut.

The six bands, in Korean

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010006.21 19.4%
101–50040065.57 17.3%
501–1,000500125.16 7.5%
1,001–2,0001,000304.85 7.3%
2,001–4,0002,000804.54 7.2%
4,001–8,0004,0001644.21 6.9%

Word shape

Character lengths across the Korean top 1,000 — average 2.5.

1
2
3
4
5
6
7
9

Characters per word

Phrases that behave like words

381 multi-word units earned a rank of their own.

  • #129 안 돼 No
  • #329 것 같아 That it happened / I think so.
  • #421 안 돼요 No
  • #423 것 같아요 I think so
  • #437 i don't I don't
  • #444 you know You know
  • #534 더 이상 Any more / No more
  • #615 둘 다 Both

What stands out in Korean

Read straight off this language's own numbers.

  • Just 1,738 entries cover half of everything said in the Korean subtitle corpus.
  • The ten most common Korean entries alone account for 4.5% of running text.
  • 381 of the top 9,026 entries are multi-word phrases that behave like single units — the most common is “안 돼”.
  • The longest single word inside the Korean top 1,000 is “something” (9 characters, rank 635).
  • Most of the Korean top 1,000 is short: 2-character words are the single largest group, and the average is 2.5 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 6.9% of the corpus.

How this Korean ranking was calculated

Built once, offline, and shipped as static data.

The source is the Korean side of OpenSubtitles v2018 via OPUS — 8,360,774 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Korean's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 722,759 distinct words appear in all; the 9,026 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 381 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 191 of the 9,026 Korean entries are still unresolved and are never used as game questions.

Common questions about Korean word frequency

Answered from this language's own corpus.

What are the most common Korean words?

The ten most common Korean entries are 내가 (I), 난 (I), 그 (That), 안 (No), 내 (My / Mine), 수 (Number / Thursday), 이 (This), 네 (Yes), 거야 (That / That's it.) and you (you). Together they account for 4.5% of all running Korean text in the corpus.

What is the most common word in Korean?

The most common Korean word is “내가”, meaning “I”. It appears 45,864 times across the corpus, which is 0.5% of everything said.

How is this Korean frequency list calculated?

The list is built from the Korean side of the OpenSubtitles corpus — 8,360,774 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 722,759 distinct forms found are ranked by raw count. The top 9,026 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Korean list include phrases as well as words?

Yes. 381 of the 9,026 Korean entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “안 돼” (No), at rank 129. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Korean do the top 100 words cover?

The 100 most common Korean entries cover 19.4% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Korean frequency dictionary free to download?

Yes. All 8,835 translated Korean entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Korean's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm