Romanian frequency dictionary

The Romanian frequency dictionary ranks the 9,708 most common Romanian words and phrases, measured across 14,922,881 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 118 Romanian entries covers half of everything said, and the top 3,866 covers 80%. The single most common Romanian word is “nu” (no), which accounts for 2.9% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 14.9M of speech
Distinct 215.5K word forms
Phrases 2,277 promoted
Glossed 9,345 to English
84.5% covered

These 9,708 entries account for 84.5% of all running Romanian text in the corpus.

118 of them cover half of it.

Romanian at a glance

The short answers, straight from the corpus.

Words for half of speech
118

Entries needed to cover 50% of everything said in Romanian

Words for 80%
3,866

Entries needed to cover 80% of running Romanian text

Most common word
nu

“no” — said 431,771 times in the Romanian corpus

Corpus size
14,922,881 tokens

215,545 distinct Romanian word forms were counted

Game modes

Mixed rounds

Romanian · All three modes, shuffled

Loading Romanian…

The ten most common Romanian entries

Together they cover 18.0% of everything said in the corpus.

  1. 1 nu no 431.8K×
  2. 2 de by / from 410.2K×
  3. 3 to 369.8K×
  4. 4 o a 264.8K×
  5. 5 a a 225.5K×
  6. 6 e is 216.1K×
  7. 7 ce what 212.3K×
  8. 8 am I have 202.5K×
  9. 9 în in 180.8K×
  10. 10 la at 175.4K×

How far each slice of the list gets you

Share of running Romanian text covered by everything up to that rank.

118 entries reach half of all Romanian speech.
3,866 entries reach 80% of all Romanian speech.
Even all 9,708 entries stop short of 90% of Romanian speech — the remainder is spread across the long tail below the cut.

The six bands, in Romanian

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010046.40 48.2%
101–500400755.51 16.4%
501–1,0005001255.03 5.7%
1,001–2,0001,0002624.69 5.2%
2,001–4,0002,0004674.35 4.7%
4,001–8,0004,0009114.00 4.3%

Word shape

Character lengths across the Romanian top 1,000 — average 4.8.

1
2
3
4
5
6
7
8
9
10
11
13

Characters per word

Phrases that behave like words

2,277 multi-word units earned a rank of their own.

  • #50 a fost was
  • #51 trebuie să must
  • #73 s a to
  • #92 cred că I think
  • #103 ar fi would be
  • #107 vreau să I want to
  • #124 vrei să do you want to
  • #125 nevoie de need / need for

What stands out in Romanian

Read straight off this language's own numbers.

  • Just 118 entries cover half of everything said in the Romanian subtitle corpus.
  • The ten most common Romanian entries alone account for 18.0% of running text.
  • 2,277 of the top 9,708 entries are multi-word phrases that behave like single units — the most common is “a fost”.
  • The longest single word inside the Romanian top 1,000 is “dumneavoastră” (13 characters, rank 817).
  • Most of the Romanian top 1,000 is short: 4-character words are the single largest group, and the average is 4.8 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.3% of the corpus.

How this Romanian ranking was calculated

Built once, offline, and shipped as static data.

The source is the Romanian side of OpenSubtitles v2024 via OPUS — 14,922,881 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Romanian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 215,545 distinct words appear in all; the 9,708 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 2,277 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 363 of the 9,708 Romanian entries are still unresolved and are never used as game questions.

Common questions about Romanian word frequency

Answered from this language's own corpus.

How many Romanian words do you need to know?

Around 118 Romanian words and phrases cover half of everything said in ordinary speech, and about 3,866 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,708 reaches 84.5%, so the last few thousand entries add far less than the first few hundred.

What are the most common Romanian words?

The ten most common Romanian entries are nu (no), de (by / from), să (to), o (a), a (a), e (is), ce (what), am (I have), în (in) and la (at). Together they account for 18.0% of all running Romanian text in the corpus.

What is the most common word in Romanian?

The most common Romanian word is “nu”, meaning “no”. It appears 431,771 times across the corpus, which is 2.9% of everything said.

How is this Romanian frequency list calculated?

The list is built from the Romanian side of the OpenSubtitles corpus — 14,922,881 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 215,545 distinct forms found are ranked by raw count. The top 9,708 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Romanian list include phrases as well as words?

Yes. 2,277 of the 9,708 Romanian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “a fost” (was), at rank 50. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Romanian do the top 100 words cover?

The 100 most common Romanian entries cover 48.2% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Romanian frequency dictionary free to download?

Yes. All 9,345 translated Romanian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Romanian's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm