Japanese frequency dictionary

The Japanese frequency dictionary ranks the 8,694 most common Japanese words and phrases, measured across 18,956,894 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 90 Japanese entries covers half of everything said, and the top 2,668 covers 80%. The single most common Japanese word is “の” (no / of), which accounts for 4.3% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 19M of speech
Distinct 92.8K word forms
Phrases 2,713 promoted
Glossed 8,660 to English
87.2% covered

These 8,694 entries account for 87.2% of all running Japanese text in the corpus.

90 of them cover half of it.

Japanese at a glance

The short answers, straight from the corpus.

Words for half of speech
90

Entries needed to cover 50% of everything said in Japanese

Words for 80%
2,668

Entries needed to cover 80% of running Japanese text

Most common word

“no / of” — said 822,587 times in the Japanese corpus

Corpus size
18,956,894 tokens

92,774 distinct Japanese word forms were counted

Game modes

Mixed rounds

Japanese · All three modes, shuffled

Loading Japanese…

The ten most common Japanese entries

Together they cover 24.3% of everything said in the corpus.

  1. 1 no / of 822.6K×
  2. 2 wa / is 689.6K×
  3. 3 o / to 567.2K×
  4. 4 ni / to 556.6K×
  5. 5 da / Yeah. 467.5K×
  6. 6 ga / that 462.2K×
  7. 7 te 286.9K×
  8. 8 ない nai / No 269.5K×
  9. 9 って tte / That's it. 249.8K×
  10. 10 ka 232.5K×

How far each slice of the list gets you

Share of running Japanese text covered by everything up to that rank.

90 entries reach half of all Japanese speech.
2,668 entries reach 80% of all Japanese speech.
Even all 8,694 entries stop short of 90% of Japanese speech — the remainder is spread across the long tail below the cut.

The six bands, in Japanese

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 100100156.33 51.0%
101–5004001315.49 15.0%
501–1,0005001505.06 6.0%
1,001–2,0001,0002804.74 5.8%
2,001–4,0002,0005704.40 5.3%
4,001–8,0004,0001,3274.00 4.2%

Word shape

Character lengths across the Japanese top 1,000 — average 2.1.

1
2
3
4
5
6
8

Characters per word

Phrases that behave like words

2,713 multi-word units earned a rank of their own.

  • #15 っ た tta / It was
  • #31 ん だ is / Hmm.
  • #42 って る is / That's it.
  • #61 じゃ ない isn’t / Not really.
  • #67 て いる te iru / It is
  • #68 てく れ please
  • #69 が ある there is
  • #75 なか っ た wasn’t / I didn't.

What stands out in Japanese

Read straight off this language's own numbers.

  • Just 90 entries cover half of everything said in the Japanese subtitle corpus.
  • The ten most common Japanese entries alone account for 24.3% of running text.
  • 2,713 of the top 8,694 entries are multi-word phrases that behave like single units — the most common is “っ た”.
  • The longest single word inside the Japanese top 1,000 is “ch00ffff” (8 characters, rank 560).
  • Most of the Japanese top 1,000 is short: 2-character words are the single largest group, and the average is 2.1 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.2% of the corpus.

How this Japanese ranking was calculated

Built once, offline, and shipped as static data.

The source is the Japanese side of OpenSubtitles v2018 via OPUS — 18,956,894 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Japanese's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 92,774 distinct words appear in all; the 8,694 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 2,713 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 34 of the 8,694 Japanese entries are still unresolved and are never used as game questions.

Common questions about Japanese word frequency

Answered from this language's own corpus.

How many Japanese words do you need to know?

Around 90 Japanese words and phrases cover half of everything said in ordinary speech, and about 2,668 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 8,694 reaches 87.2%, so the last few thousand entries add far less than the first few hundred.

What are the most common Japanese words?

The ten most common Japanese entries are の (no / of), は (wa / is), を (o / to), に (ni / to), だ (da / Yeah.), が (ga / that), て (te), ない (nai / No), って (tte / That's it.) and か (ka). Together they account for 24.3% of all running Japanese text in the corpus.

What is the most common word in Japanese?

The most common Japanese word is “の”, meaning “no / of”. It appears 822,587 times across the corpus, which is 4.3% of everything said.

How is this Japanese frequency list calculated?

The list is built from the Japanese side of the OpenSubtitles corpus — 18,956,894 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 92,774 distinct forms found are ranked by raw count. The top 8,694 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Japanese list include phrases as well as words?

Yes. 2,713 of the 8,694 Japanese entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “っ た” (tta / It was), at rank 15. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Japanese do the top 100 words cover?

The 100 most common Japanese entries cover 51.0% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Japanese frequency dictionary free to download?

Yes. All 8,660 translated Japanese entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Japanese's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm