English
frequency dictionary
The English frequency dictionary ranks the 9,776 most common English words and phrases, measured across 17,193,396 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 84 English entries covers half of everything said, and the top 1,235 covers 80%. The single most common English word is “you” (you), which accounts for 3.7% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,776 entries account for 91.1% of all running English text in the corpus.
84 of them cover half of it.
English at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 84
- Words for 80%
- 1,235
- Most common word
- you
- Corpus size
- 17,193,396 tokens
Entries needed to cover 50% of everything said in English
Entries needed to cover 80% of running English text
“you” — said 634,603 times in the English corpus
135,757 distinct English word forms were counted
Mixed rounds
English · All three modes, shuffled
Loading English…
The ten most common English entries
Together they cover 19.7% of everything said in the corpus.
- 1 you you 634.6K×
- 2 the the 529.4K×
- 3 i I 479.2K×
- 4 to to 382.8K×
- 5 a a 333.5K×
- 6 and and 236.4K×
- 7 it it 230.8K×
- 8 of of 196.9K×
- 9 that that 191.6K×
- 10 is is 180.1K×
How far each slice of the list gets you
Share of running English text covered by everything up to that rank.
The six bands, in English
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 2 | 6.51 | 52.7% |
| 101–500 | 400 | 52 | 5.58 | 19.4% |
| 501–1,000 | 500 | 77 | 5.05 | 6.2% |
| 1,001–2,000 | 1,000 | 121 | 4.69 | 5.1% |
| 2,001–4,000 | 2,000 | 249 | 4.30 | 4.3% |
| 4,001–8,000 | 4,000 | 463 | 3.90 | 3.4% |
Word shape
Character lengths across the English top 1,000 — average 4.9.
Characters per word
Phrases that behave like words
1,142 multi-word units earned a rank of their own.
- #72 you know You know
- #85 come on Come on
- #113 i know I know
- #119 thank you Thank you
- #124 want to want to
- #134 going to going to
- #139 i think I think
- #153 you want you want
What stands out in English
Read straight off this language's own numbers.
- Just 84 entries cover half of everything said in the English subtitle corpus.
- The ten most common English entries alone account for 19.7% of running text.
- 1,142 of the top 9,776 entries are multi-word phrases that behave like single units — the most common is “you know”.
- The longest single word inside the English top 1,000 is “motherfucker” (12 characters, rank 948).
- Most of the English top 1,000 is short: 4-character words are the single largest group, and the average is 4.9 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 3.4% of the corpus.
How this English ranking was calculated
Built once, offline, and shipped as static data.
The source is the English side of OpenSubtitles v2024 via OPUS — 17,193,396 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using English's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 135,757 distinct words appear in all; the 9,776 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,142 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation.
Common questions about English word frequency
Answered from this language's own corpus.
How many English words do you need to know?
Around 84 English words and phrases cover half of everything said in ordinary speech, and about 1,235 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,776 reaches 91.1%, so the last few thousand entries add far less than the first few hundred.
What are the most common English words?
The ten most common English entries are you (you), the (the), i (I), to (to), a (a), and (and), it (it), of (of), that (that) and is (is). Together they account for 19.7% of all running English text in the corpus.
What is the most common word in English?
The most common English word is “you”, meaning “you”. It appears 634,603 times across the corpus, which is 3.7% of everything said.
How is this English frequency list calculated?
The list is built from the English side of the OpenSubtitles corpus — 17,193,396 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 135,757 distinct forms found are ranked by raw count. The top 9,776 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the English list include phrases as well as words?
Yes. 1,142 of the 9,776 English entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “you know” (You know), at rank 72. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much English do the top 100 words cover?
The 100 most common English entries cover 52.7% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the English frequency dictionary free to download?
Yes. All 9,776 translated English entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to English's.