Catalan frequency dictionary

The Catalan frequency dictionary ranks the 9,072 most common Catalan words and phrases, measured across 3,471,714 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 100 Catalan entries covers half of everything said, and the top 2,561 covers 80%. The single most common Catalan word is “que” (that), which accounts for 3.1% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 3.5M of speech
Distinct 84.5K word forms
Phrases 2,085 promoted
Glossed 8,808 to English
87.2% covered

These 9,072 entries account for 87.2% of all running Catalan text in the corpus.

100 of them cover half of it.

Catalan at a glance

The short answers, straight from the corpus.

Words for half of speech
100

Entries needed to cover 50% of everything said in Catalan

Words for 80%
2,561

Entries needed to cover 80% of running Catalan text

Most common word
que

“that” — said 107,880 times in the Catalan corpus

Corpus size
3,471,714 tokens

84,493 distinct Catalan word forms were counted

Game modes

Mixed rounds

Catalan · All three modes, shuffled

Loading Catalan…

The ten most common Catalan entries

Together they cover 21.9% of everything said in the corpus.

  1. 1 que that 107.9K×
  2. 2 no no 100.8K×
  3. 3 de of / from 96.2K×
  4. 4 la the 94.6K×
  5. 5 el the 78K×
  6. 6 a to 76.9K×
  7. 7 i and 61.4K×
  8. 8 és is 53.9K×
  9. 9 per for / by 46.6K×
  10. 10 un a / one 45K×

How far each slice of the list gets you

Share of running Catalan text covered by everything up to that rank.

100 entries reach half of all Catalan speech.
2,561 entries reach 80% of all Catalan speech.
Even all 9,072 entries stop short of 90% of Catalan speech — the remainder is spread across the long tail below the cut.

The six bands, in Catalan

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010076.38 50.1%
101–500400645.54 17.3%
501–1,000500975.03 5.7%
1,001–2,0001,0002054.68 5.1%
2,001–4,0002,0004984.35 4.7%
4,001–8,0004,0009554.00 4.2%

Word shape

Character lengths across the Catalan top 1,000 — average 5.2.

1
2
3
4
5
6
7
8
9
10
11
12

Characters per word

Phrases that behave like words

2,085 multi-word units earned a rank of their own.

  • #52 hi ha there is / There are
  • #59 la meva my / mine
  • #63 el meu my / mine
  • #77 la seva his / his/her
  • #89 la teva your / yours
  • #91 el seu his
  • #97 el teu your / yours
  • #106 no hi not there

What stands out in Catalan

Read straight off this language's own numbers.

  • Just 100 entries cover half of everything said in the Catalan subtitle corpus.
  • The ten most common Catalan entries alone account for 21.9% of running text.
  • 2,085 of the top 9,072 entries are multi-word phrases that behave like single units — the most common is “hi ha”.
  • The longest single word inside the Catalan top 1,000 is “probablement” (12 characters, rank 552).
  • Most of the Catalan top 1,000 is short: 5-character words are the single largest group, and the average is 5.2 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.2% of the corpus.

How this Catalan ranking was calculated

Built once, offline, and shipped as static data.

The source is the Catalan side of OpenSubtitles v2018 via OPUS — 3,471,714 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Catalan's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 84,493 distinct words appear in all; the 9,072 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 2,085 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 264 of the 9,072 Catalan entries are still unresolved and are never used as game questions.

Common questions about Catalan word frequency

Answered from this language's own corpus.

How many Catalan words do you need to know?

Around 100 Catalan words and phrases cover half of everything said in ordinary speech, and about 2,561 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,072 reaches 87.2%, so the last few thousand entries add far less than the first few hundred.

What are the most common Catalan words?

The ten most common Catalan entries are que (that), no (no), de (of / from), la (the), el (the), a (to), i (and), és (is), per (for / by) and un (a / one). Together they account for 21.9% of all running Catalan text in the corpus.

What is the most common word in Catalan?

The most common Catalan word is “que”, meaning “that”. It appears 107,880 times across the corpus, which is 3.1% of everything said.

How is this Catalan frequency list calculated?

The list is built from the Catalan side of the OpenSubtitles corpus — 3,471,714 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 84,493 distinct forms found are ranked by raw count. The top 9,072 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Catalan list include phrases as well as words?

Yes. 2,085 of the 9,072 Catalan entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “hi ha” (there is / There are), at rank 52. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Catalan do the top 100 words cover?

The 100 most common Catalan entries cover 50.1% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Catalan frequency dictionary free to download?

Yes. All 8,808 translated Catalan entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Catalan's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm