Slovenian frequency dictionary

The Slovenian frequency dictionary ranks the 9,270 most common Slovenian words and phrases, measured across 267,368,007 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 171 Slovenian entries covers half of everything said, and the top 4,791 covers 80%. The single most common Slovenian word is “je” (is), which accounts for 3.9% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.

Corpus

Tokens 267.4M of speech
Distinct 952.3K word forms
Phrases 938 promoted
Glossed 9,204 to English
83.5% covered

These 9,270 entries account for 83.5% of all running Slovenian text in the corpus.

171 of them cover half of it.

Slovenian at a glance

The short answers, straight from the corpus.

Words for half of speech
171

Entries needed to cover 50% of everything said in Slovenian

Words for 80%
4,791

Entries needed to cover 80% of running Slovenian text

Most common word
je

“is” — said 10,547,795 times in the Slovenian corpus

Corpus size
267,368,007 tokens

952,295 distinct Slovenian word forms were counted

Game modes

Mixed rounds

Slovenian · All three modes, shuffled

Loading Slovenian…

The ten most common Slovenian entries

Together they cover 18.2% of everything said in the corpus.

  1. 1 je is 10.5M×
  2. 2 ne no 6.2M×
  3. 3 da yes 5.4M×
  4. 4 se se 5.2M×
  5. 5 v in 4.1M×
  6. 6 sem I am 3.8M×
  7. 7 to to 3.7M×
  8. 8 in and 3.6M×
  9. 9 si si 3.1M×
  10. 10 kaj what 3M×

How far each slice of the list gets you

Share of running Slovenian text covered by everything up to that rank.

171 entries reach half of all Slovenian speech.
4,791 entries reach 80% of all Slovenian speech.
Even all 9,270 entries stop short of 90% of Slovenian speech — the remainder is spread across the long tail below the cut.

The six bands, in Slovenian

The same bands the game asks you to guess.

BandEntriesPhrasesMedian ZipfText covered
Top 10010036.38 44.2%
101–500400215.54 16.9%
501–1,000500425.08 6.4%
1,001–2,0001,000944.74 5.8%
2,001–4,0002,0002074.40 5.3%
4,001–8,0004,0004124.06 4.8%

Word shape

Character lengths across the Slovenian top 1,000 — average 5.1.

1
2
3
4
5
6
7
8
9
10
11
12

Characters per word

Phrases that behave like words

938 multi-word units earned a rank of their own.

  • #51 v redu okay
  • #95 ne vem I don't know
  • #97 je bil was
  • #102 mislim da I think
  • #116 je bilo it was / was
  • #118 ne morem I can't
  • #159 kje je where is it
  • #170 še vedno still

What stands out in Slovenian

Read straight off this language's own numbers.

  • Just 171 entries cover half of everything said in the Slovenian subtitle corpus.
  • The ten most common Slovenian entries alone account for 18.2% of running text.
  • 938 of the top 9,270 entries are multi-word phrases that behave like single units — the most common is “v redu”.
  • The longest single word inside the Slovenian top 1,000 is “pripravljeni” (12 characters, rank 637).
  • Most of the Slovenian top 1,000 is short: 5-character words are the single largest group, and the average is 5.1 characters.
  • Frequency falls away fast: the whole 4,001–8,000 band together covers only 4.8% of the corpus.

How this Slovenian ranking was calculated

Built once, offline, and shipped as static data.

The source is the Slovenian side of OpenSubtitles v2018 via OPUS — 267,368,007 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.

Words are counted using Slovenian's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 952,295 distinct words appear in all; the 9,270 most common are kept.

Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 938 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.

English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 66 of the 9,270 Slovenian entries are still unresolved and are never used as game questions.

Common questions about Slovenian word frequency

Answered from this language's own corpus.

How many Slovenian words do you need to know?

Around 171 Slovenian words and phrases cover half of everything said in ordinary speech, and about 4,791 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,270 reaches 83.5%, so the last few thousand entries add far less than the first few hundred.

What are the most common Slovenian words?

The ten most common Slovenian entries are je (is), ne (no), da (yes), se (se), v (in), sem (I am), to (to), in (and), si (si) and kaj (what). Together they account for 18.2% of all running Slovenian text in the corpus.

What is the most common word in Slovenian?

The most common Slovenian word is “je”, meaning “is”. It appears 10,547,795 times across the corpus, which is 3.9% of everything said.

How is this Slovenian frequency list calculated?

The list is built from the Slovenian side of the OpenSubtitles corpus — 267,368,007 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 952,295 distinct forms found are ranked by raw count. The top 9,270 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.

Does the Slovenian list include phrases as well as words?

Yes. 938 of the 9,270 Slovenian entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “v redu” (okay), at rank 51. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.

How much Slovenian do the top 100 words cover?

The 100 most common Slovenian entries cover 44.2% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.

Is the Slovenian frequency dictionary free to download?

Yes. All 9,204 translated Slovenian entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.

Corpora of a similar size

Languages whose subtitle corpus is closest to Slovenian's.

Compare side by side

KataRank — prebuilt frequency dictionaries. No account, no scraping, no translation calls.

© 2026 KataRank.com Made with love in Stockholm