Korean
frequency dictionary
The Korean frequency dictionary ranks the 9,026 most common Korean words and phrases, measured across 8,360,774 tokens of subtitle dialogue from the OpenSubtitles corpus. The single most common Korean word is “내가” (I), which accounts for 0.5% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,026 entries account for 65.5% of all running Korean text in the corpus.
1,738 of them cover half of it.
Korean at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 1,738
- Most common word
- 내가
- Corpus size
- 8,360,774 tokens
Entries needed to cover 50% of everything said in Korean
“I” — said 45,864 times in the Korean corpus
722,759 distinct Korean word forms were counted
Mixed rounds
Korean · All three modes, shuffled
Loading Korean…
The ten most common Korean entries
Together they cover 4.5% of everything said in the corpus.
- 1 내가 I 45.9K×
- 2 난 I 41.1K×
- 3 그 That 40.9K×
- 4 안 No 39.7K×
- 5 내 My / Mine 39.2K×
- 6 수 Number / Thursday 38.6K×
- 7 이 This 35.8K×
- 8 네 Yes 34.2K×
- 9 거야 That / That's it. 33.5K×
- 10 you you 29.9K×
How far each slice of the list gets you
Share of running Korean text covered by everything up to that rank.
The six bands, in Korean
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 0 | 6.21 | 19.4% |
| 101–500 | 400 | 6 | 5.57 | 17.3% |
| 501–1,000 | 500 | 12 | 5.16 | 7.5% |
| 1,001–2,000 | 1,000 | 30 | 4.85 | 7.3% |
| 2,001–4,000 | 2,000 | 80 | 4.54 | 7.2% |
| 4,001–8,000 | 4,000 | 164 | 4.21 | 6.9% |
Word shape
Character lengths across the Korean top 1,000 — average 2.5.
Characters per word
Phrases that behave like words
381 multi-word units earned a rank of their own.
- #129 안 돼 No
- #329 것 같아 That it happened / I think so.
- #421 안 돼요 No
- #423 것 같아요 I think so
- #437 i don't I don't
- #444 you know You know
- #534 더 이상 Any more / No more
- #615 둘 다 Both
What stands out in Korean
Read straight off this language's own numbers.
- Just 1,738 entries cover half of everything said in the Korean subtitle corpus.
- The ten most common Korean entries alone account for 4.5% of running text.
- 381 of the top 9,026 entries are multi-word phrases that behave like single units — the most common is “안 돼”.
- The longest single word inside the Korean top 1,000 is “something” (9 characters, rank 635).
- Most of the Korean top 1,000 is short: 2-character words are the single largest group, and the average is 2.5 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 6.9% of the corpus.
How this Korean ranking was calculated
Built once, offline, and shipped as static data.
The source is the Korean side of OpenSubtitles v2018 via OPUS — 8,360,774 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Korean's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 722,759 distinct words appear in all; the 9,026 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 381 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 191 of the 9,026 Korean entries are still unresolved and are never used as game questions.
Common questions about Korean word frequency
Answered from this language's own corpus.
What are the most common Korean words?
The ten most common Korean entries are 내가 (I), 난 (I), 그 (That), 안 (No), 내 (My / Mine), 수 (Number / Thursday), 이 (This), 네 (Yes), 거야 (That / That's it.) and you (you). Together they account for 4.5% of all running Korean text in the corpus.
What is the most common word in Korean?
The most common Korean word is “내가”, meaning “I”. It appears 45,864 times across the corpus, which is 0.5% of everything said.
How is this Korean frequency list calculated?
The list is built from the Korean side of the OpenSubtitles corpus — 8,360,774 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 722,759 distinct forms found are ranked by raw count. The top 9,026 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Korean list include phrases as well as words?
Yes. 381 of the 9,026 Korean entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “안 돼” (No), at rank 129. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Korean do the top 100 words cover?
The 100 most common Korean entries cover 19.4% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Korean frequency dictionary free to download?
Yes. All 8,835 translated Korean entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Korean's.