Chinese–English subtitles (English side)
frequency dictionary
The Chinese–English subtitles (English side) frequency dictionary ranks the 9,429 most common Chinese–English subtitles (English side) words and phrases, measured across 46,578,182 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 116 Chinese–English subtitles (English side) entries covers half of everything said, and the top 3,580 covers 80%. The single most common Chinese–English subtitles (English side) word is “you” (you), which accounts for 3.1% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,429 entries account for 83.9% of all running Chinese–English subtitles (English side) text in the corpus.
116 of them cover half of it.
Chinese–English subtitles (English side) at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 116
- Words for 80%
- 3,580
- Most common word
- you
- Corpus size
- 46,578,182 tokens
Entries needed to cover 50% of everything said in Chinese–English subtitles (English side)
Entries needed to cover 80% of running Chinese–English subtitles (English side) text
“you” — said 1,455,535 times in the Chinese–English subtitles (English side) corpus
231,139 distinct Chinese–English subtitles (English side) word forms were counted
Mixed rounds
Chinese–English subtitles (English side) · All three modes, shuffled
Loading Chinese–English subtitles (English side)…
The ten most common Chinese–English subtitles (English side) entries
Together they cover 18.4% of everything said in the corpus.
- 1 you you 1.5M×
- 2 the the 1.3M×
- 3 i I 1.1M×
- 4 to to 958K×
- 5 a a 825.7K×
- 6 chffffff 3ch2f2f2f 4ch000000 chffffff 3ch2f2f2f 4ch000000 705.3K×
- 7 chffffff 3ch000000 4ch000000 chffffff 3ch000000 4ch000000 612.8K×
- 8 and and 583.8K×
- 9 it it 522.5K×
- 10 of of 504.5K×
How far each slice of the list gets you
Share of running Chinese–English subtitles (English side) text covered by everything up to that rank.
The six bands, in Chinese–English subtitles (English side)
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 5 | 6.46 | 48.1% |
| 101–500 | 400 | 72 | 5.54 | 17.9% |
| 501–1,000 | 500 | 102 | 5.03 | 5.7% |
| 1,001–2,000 | 1,000 | 160 | 4.66 | 4.8% |
| 2,001–4,000 | 2,000 | 339 | 4.29 | 4.1% |
| 4,001–8,000 | 4,000 | 443 | 3.90 | 3.3% |
Word shape
Character lengths across the Chinese–English subtitles (English side) top 1,000 — average 4.7.
Characters per word
Phrases that behave like words
1,288 multi-word units earned a rank of their own.
- #6 chffffff 3ch2f2f2f 4ch000000 chffffff 3ch2f2f2f 4ch000000
- #7 chffffff 3ch000000 4ch000000 chffffff 3ch000000 4ch000000
- #60 chffffff 3ch111111 4ch111111 chffffff 3ch111111 4ch111111
- #68 i don't I don't
- #78 you know You know
- #110 come on Come on
- #119 want to want to
- #121 i know I know
What stands out in Chinese–English subtitles (English side)
Read straight off this language's own numbers.
- Just 116 entries cover half of everything said in the Chinese–English subtitles (English side) subtitle corpus.
- The ten most common Chinese–English subtitles (English side) entries alone account for 18.4% of running text.
- 1,288 of the top 9,429 entries are multi-word phrases that behave like single units — the most common is “chffffff 3ch2f2f2f 4ch000000”.
- The longest single word inside the Chinese–English subtitles (English side) top 1,000 is “information” (11 characters, rank 867).
- Most of the Chinese–English subtitles (English side) top 1,000 is short: 4-character words are the single largest group, and the average is 4.7 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 3.3% of the corpus.
How this Chinese–English subtitles (English side) ranking was calculated
Built once, offline, and shipped as static data.
The source is the Chinese–English subtitles (English side) side of OpenSubtitles v2018 via OPUS — 46,578,182 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Chinese–English subtitles (English side)'s own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 231,139 distinct words appear in all; the 9,429 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,288 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation.
Common questions about Chinese–English subtitles (English side) word frequency
Answered from this language's own corpus.
How many Chinese–English subtitles (English side) words do you need to know?
Around 116 Chinese–English subtitles (English side) words and phrases cover half of everything said in ordinary speech, and about 3,580 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,429 reaches 83.9%, so the last few thousand entries add far less than the first few hundred.
What are the most common Chinese–English subtitles (English side) words?
The ten most common Chinese–English subtitles (English side) entries are you (you), the (the), i (I), to (to), a (a), chffffff 3ch2f2f2f 4ch000000 (chffffff 3ch2f2f2f 4ch000000), chffffff 3ch000000 4ch000000 (chffffff 3ch000000 4ch000000), and (and), it (it) and of (of). Together they account for 18.4% of all running Chinese–English subtitles (English side) text in the corpus.
What is the most common word in Chinese–English subtitles (English side)?
The most common Chinese–English subtitles (English side) word is “you”, meaning “you”. It appears 1,455,535 times across the corpus, which is 3.1% of everything said.
How is this Chinese–English subtitles (English side) frequency list calculated?
The list is built from the Chinese–English subtitles (English side) side of the OpenSubtitles corpus — 46,578,182 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 231,139 distinct forms found are ranked by raw count. The top 9,429 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Chinese–English subtitles (English side) list include phrases as well as words?
Yes. 1,288 of the 9,429 Chinese–English subtitles (English side) entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “chffffff 3ch2f2f2f 4ch000000” (chffffff 3ch2f2f2f 4ch000000), at rank 6. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Chinese–English subtitles (English side) do the top 100 words cover?
The 100 most common Chinese–English subtitles (English side) entries cover 48.1% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Chinese–English subtitles (English side) frequency dictionary free to download?
Yes. All 9,429 translated Chinese–English subtitles (English side) entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Chinese–English subtitles (English side)'s.