Spanish
frequency dictionary
The Spanish frequency dictionary ranks the 9,625 most common Spanish words and phrases, measured across 14,107,392 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 101 Spanish entries covers half of everything said, and the top 2,891 covers 80%. The single most common Spanish word is “que” (that), which accounts for 3.2% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,625 entries account for 86.4% of all running Spanish text in the corpus.
101 of them cover half of it.
Spanish at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 101
- Words for 80%
- 2,891
- Most common word
- que
- Corpus size
- 14,107,392 tokens
Entries needed to cover 50% of everything said in Spanish
Entries needed to cover 80% of running Spanish text
“that” — said 449,148 times in the Spanish corpus
195,216 distinct Spanish word forms were counted
Mixed rounds
Spanish · All three modes, shuffled
Loading Spanish…
The ten most common Spanish entries
Together they cover 21.8% of everything said in the corpus.
- 1 que that 449.1K×
- 2 de from 445K×
- 3 no no 424.9K×
- 4 a to 318.6K×
- 5 la the 296.3K×
- 6 el the 251.6K×
- 7 es is 241.1K×
- 8 y and 227.5K×
- 9 en in 216.9K×
- 10 lo it 203K×
How far each slice of the list gets you
Share of running Spanish text covered by everything up to that rank.
The six bands, in Spanish
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 1 | 6.34 | 49.9% |
| 101–500 | 400 | 50 | 5.49 | 16.5% |
| 501–1,000 | 500 | 92 | 5.05 | 5.8% |
| 1,001–2,000 | 1,000 | 209 | 4.69 | 5.1% |
| 2,001–4,000 | 2,000 | 482 | 4.36 | 4.8% |
| 4,001–8,000 | 4,000 | 966 | 4.01 | 4.3% |
Word shape
Character lengths across the Spanish top 1,000 — average 5.4.
Characters per word
Phrases that behave like words
2,233 multi-word units earned a rank of their own.
- #88 por favor please
- #101 voy a I'm going to
- #102 creo que I think
- #127 lo siento I'm sorry
- #137 va a is going to
- #139 no puedo I can't
- #159 de acuerdo okay / Agreed
- #165 un poco a little
What stands out in Spanish
Read straight off this language's own numbers.
- Just 101 entries cover half of everything said in the Spanish subtitle corpus.
- The ten most common Spanish entries alone account for 21.8% of running text.
- 2,233 of the top 9,625 entries are multi-word phrases that behave like single units — the most common is “por favor”.
- The longest single word inside the Spanish top 1,000 is “probablemente” (13 characters, rank 648).
- Most of the Spanish top 1,000 is short: 5-character words are the single largest group, and the average is 5.4 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.3% of the corpus.
How this Spanish ranking was calculated
Built once, offline, and shipped as static data.
The source is the Spanish side of OpenSubtitles v2024 via OPUS — 14,107,392 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Spanish's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 195,216 distinct words appear in all; the 9,625 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 2,233 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 221 of the 9,625 Spanish entries are still unresolved and are never used as game questions.
Common questions about Spanish word frequency
Answered from this language's own corpus.
How many Spanish words do you need to know?
Around 101 Spanish words and phrases cover half of everything said in ordinary speech, and about 2,891 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,625 reaches 86.4%, so the last few thousand entries add far less than the first few hundred.
What are the most common Spanish words?
The ten most common Spanish entries are que (that), de (from), no (no), a (to), la (the), el (the), es (is), y (and), en (in) and lo (it). Together they account for 21.8% of all running Spanish text in the corpus.
What is the most common word in Spanish?
The most common Spanish word is “que”, meaning “that”. It appears 449,148 times across the corpus, which is 3.2% of everything said.
How is this Spanish frequency list calculated?
The list is built from the Spanish side of the OpenSubtitles corpus — 14,107,392 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 195,216 distinct forms found are ranked by raw count. The top 9,625 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Spanish list include phrases as well as words?
Yes. 2,233 of the 9,625 Spanish entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “por favor” (please), at rank 88. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Spanish do the top 100 words cover?
The 100 most common Spanish entries cover 49.9% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Spanish frequency dictionary free to download?
Yes. All 9,404 translated Spanish entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Spanish's.