Vietnamese frequency dictionary

The Vietnamese frequency dictionary ranks the 22,000 most common Vietnamese words and phrases, measured across 19,798,574 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 91 Vietnamese entries covers half of everything said, and the top 618 covers 80%. The single most common Vietnamese word is “tôi” (I), which accounts for 2.7% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 19.8M of speech
Distinct 77.7K word forms
Phrases 2,481 promoted
Glossed 10,538 to English
93.3% covered

These 22,000 entries account for 93.3% of all running Vietnamese text in the corpus.

91 of them cover half of it.

Vietnamese at a glance

The short answers, straight from the corpus.

Words for half of speech
91

Entries needed to cover 50% of everything said in Vietnamese

Words for 80%
618

Entries needed to cover 80% of running Vietnamese text

Most common word
tôi

“I” — said 539,682 times in the Vietnamese corpus

Corpus size
19,798,574 tokens

77,705 distinct Vietnamese word forms were counted

Game modes

Mixed rounds

Vietnamese · All three modes, shuffled

Loading Vietnamese…

The ten most common Vietnamese entries

Together they cover 15.3% of everything said in the corpus.

  1. 1 tôi I 539.7K×
  2. 2 không no 489K×
  3. 3 is 346.9K×
  4. 4 anh he / brother 338.7K×
  5. 5 yes 334.7K×
  6. 6 ta we / I 226.8K×
  7. 7 đi go 206K×
  8. 8 của of 183.9K×
  9. 9 đó that 180.7K×
  10. 10 what 180.2K×

How far each slice of the list gets you

Share of running Vietnamese text covered by everything up to that rank.

91 entries reach half of all Vietnamese speech.
618 entries reach 80% of all Vietnamese speech.
2,148 entries reach 90% of all Vietnamese speech.

The six bands, in Vietnamese

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010036.55 51.9%
101–500400465.69 25.4%
501–1,0005001095.15 7.8%
1,001–2,0001,0002974.61 4.7%
2,001–4,0002,0005284.02 2.4%
4,001–8,0004,0008333.46 1.3%

Word shape

Character lengths across the Vietnamese top 1,000 — average 3.3.

1
2
3
4
5
6
7

Characters per word

Phrases that behave like words

2,481 multi-word units earned a rank of their own.

  • #45 chúng ta we
  • #52 có thể can / possible
  • #81 chúng tôi we
  • #117 không thể can’t / impossible
  • #134 xin lỗi sorry
  • #154 tất cả all
  • #157 cảm ơn thank you
  • #188 mọi người everyone

What stands out in Vietnamese

Read straight off this language's own numbers.

  • Just 91 entries cover half of everything said in the Vietnamese subtitle corpus.
  • The ten most common Vietnamese entries alone account for 15.3% of running text.
  • 2,481 of the top 22,000 entries are multi-word phrases that behave like single units — the most common is “chúng ta”.
  • The longest single word inside the Vietnamese top 1,000 is “charlie” (7 characters, rank 994).
  • Most of the Vietnamese top 1,000 is short: 3-character words are the single largest group, and the average is 3.3 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 1.3% of the corpus.

How this Vietnamese ranking was calculated

Built once, offline, and shipped as static data.

The source is the Vietnamese side of OpenSubtitles v2024 via OPUS — 19,798,574 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Vietnamese's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 77,705 distinct words appear in all; the 22,000 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 2,481 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 11,462 of the 22,000 Vietnamese entries are still unresolved and are never used as game questions.

Common questions about Vietnamese word frequency

Answered from this language's own corpus.

How many Vietnamese words do you need to know?

Around 91 Vietnamese words and phrases cover half of everything said in ordinary speech, and about 618 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 22,000 reaches 93.3%, so the last few thousand entries add far less than the first few hundred.

What are the most common Vietnamese words?

The ten most common Vietnamese entries are tôi (I), không (no), là (is), anh (he / brother), có (yes), ta (we / I), đi (go), của (of), đó (that) and gì (what). Together they account for 15.3% of all running Vietnamese text in the corpus.

What is the most common word in Vietnamese?

The most common Vietnamese word is “tôi”, meaning “I”. It appears 539,682 times across the corpus, which is 2.7% of everything said.

How is this Vietnamese frequency list calculated?

The list is built from the Vietnamese side of the OpenSubtitles corpus — 19,798,574 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 77,705 distinct forms found are ranked by raw count. The top 22,000 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Vietnamese list include phrases as well as words?

Yes. 2,481 of the 22,000 Vietnamese entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “chúng ta” (we), at rank 45. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Vietnamese do the top 100 words cover?

The 100 most common Vietnamese entries cover 51.9% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Vietnamese frequency dictionary free to download?

Yes. All 10,538 translated Vietnamese entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Vietnamese's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm