Indonesian
frequency dictionary
The Indonesian frequency dictionary ranks the 10,012 most common Indonesian words and phrases, measured across 12,852,227 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 118 Indonesian entries covers half of everything said, and the top 1,821 covers 80%. The single most common Indonesian word is “aku” (I), which accounts for 3.7% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 10,012 entries account for 90.7% of all running Indonesian text in the corpus.
118 of them cover half of it.
Indonesian at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 118
- Words for 80%
- 1,821
- Most common word
- aku
- Corpus size
- 12,852,227 tokens
Entries needed to cover 50% of everything said in Indonesian
Entries needed to cover 80% of running Indonesian text
“I” — said 474,467 times in the Indonesian corpus
143,900 distinct Indonesian word forms were counted
Mixed rounds
Indonesian · All three modes, shuffled
Loading Indonesian…
The ten most common Indonesian entries
Together they cover 19.1% of everything said in the corpus.
- 1 aku I 474.5K×
- 2 kau you 341.5K×
- 3 yang who / that 270.1K×
- 4 itu that 212.4K×
- 5 tidak doesn’t / No 204.5K×
- 6 ini this 204.3K×
- 7 di at / in 196.8K×
- 8 apa what 186.8K×
- 9 dia he / he/she 180.7K×
- 10 dan and 177.7K×
How far each slice of the list gets you
Share of running Indonesian text covered by everything up to that rank.
The six bands, in Indonesian
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 2 | 6.42 | 47.9% |
| 101–500 | 400 | 17 | 5.60 | 19.4% |
| 501–1,000 | 500 | 21 | 5.13 | 7.1% |
| 1,001–2,000 | 1,000 | 52 | 4.78 | 6.4% |
| 2,001–4,000 | 2,000 | 136 | 4.42 | 5.5% |
| 4,001–8,000 | 4,000 | 322 | 4.01 | 4.4% |
Word shape
Character lengths across the Indonesian top 1,000 — average 6.1.
Characters per word
Phrases that behave like words
724 multi-word units earned a rank of their own.
- #56 di sini here
- #66 terima kasih thank you
- #134 di mana where
- #138 di sana there
- #213 apa yang terjadi what happened
- #233 tentu saja of course
- #259 berada di be at / located in
- #303 luar biasa extraordinary
What stands out in Indonesian
Read straight off this language's own numbers.
- Just 118 entries cover half of everything said in the Indonesian subtitle corpus.
- The ten most common Indonesian entries alone account for 19.1% of running text.
- 724 of the top 10,012 entries are multi-word phrases that behave like single units — the most common is “di sini”.
- The longest single word inside the Indonesian top 1,000 is “mendapatkannya” (14 characters, rank 658).
- Most of the Indonesian top 1,000 is short: 5-character words are the single largest group, and the average is 6.1 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.4% of the corpus.
How this Indonesian ranking was calculated
Built once, offline, and shipped as static data.
The source is the Indonesian side of OpenSubtitles v2024 via OPUS — 12,852,227 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Indonesian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 143,900 distinct words appear in all; the 10,012 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 724 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 516 of the 10,012 Indonesian entries are still unresolved and are never used as game questions.
Common questions about Indonesian word frequency
Answered from this language's own corpus.
How many Indonesian words do you need to know?
Around 118 Indonesian words and phrases cover half of everything said in ordinary speech, and about 1,821 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 10,012 reaches 90.7%, so the last few thousand entries add far less than the first few hundred.
What are the most common Indonesian words?
The ten most common Indonesian entries are aku (I), kau (you), yang (who / that), itu (that), tidak (doesn’t / No), ini (this), di (at / in), apa (what), dia (he / he/she) and dan (and). Together they account for 19.1% of all running Indonesian text in the corpus.
What is the most common word in Indonesian?
The most common Indonesian word is “aku”, meaning “I”. It appears 474,467 times across the corpus, which is 3.7% of everything said.
How is this Indonesian frequency list calculated?
The list is built from the Indonesian side of the OpenSubtitles corpus — 12,852,227 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 143,900 distinct forms found are ranked by raw count. The top 10,012 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Indonesian list include phrases as well as words?
Yes. 724 of the 10,012 Indonesian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “di sini” (here), at rank 56. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Indonesian do the top 100 words cover?
The 100 most common Indonesian entries cover 47.9% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Indonesian frequency dictionary free to download?
Yes. All 9,496 translated Indonesian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Indonesian's.