Swedish
frequency dictionary
The Swedish frequency dictionary ranks the 9,901 most common Swedish words and phrases, measured across 14,067,135 tokens of subtitle dialogue from the OpenSubtitles corpus. Learning the top 61 Swedish entries covers half of everything said, and the top 1,484 covers 80%. The single most common Swedish word is “jag” (I), which accounts for 3.9% of running text on its own. Every entry carries its rank, raw count, Zipf score, and an English translation.
Corpus
These 9,901 entries account for 89.3% of all running Swedish text in the corpus.
61 of them cover half of it.
Swedish at a glance
The short answers, straight from the corpus.
- Words for half of speech
- 61
- Words for 80%
- 1,484
- Most common word
- jag
- Corpus size
- 14,067,135 tokens
Entries needed to cover 50% of everything said in Swedish
Entries needed to cover 80% of running Swedish text
“I” — said 550,087 times in the Swedish corpus
230,212 distinct Swedish word forms were counted
Mixed rounds
Swedish · All three modes, shuffled
Loading Swedish…
The ten most common Swedish entries
Together they cover 23.3% of everything said in the corpus.
- 1 jag I 550.1K×
- 2 det it / that 511.5K×
- 3 är is 469.8K×
- 4 du you 404.1K×
- 5 inte not 277.4K×
- 6 att that 276.6K×
- 7 en a / one 214.7K×
- 8 och and 205.6K×
- 9 vi we 191.4K×
- 10 har has 177.3K×
How far each slice of the list gets you
Share of running Swedish text covered by everything up to that rank.
The six bands, in Swedish
The same bands the game asks you to guess.
| Band | Entries | Phrases | Median Zipf | Text covered |
|---|---|---|---|---|
| Top 100 | 100 | 1 | 6.43 | 56.4% |
| 101–500 | 400 | 37 | 5.48 | 15.6% |
| 501–1,000 | 500 | 101 | 5.01 | 5.3% |
| 1,001–2,000 | 1,000 | 156 | 4.64 | 4.6% |
| 2,001–4,000 | 2,000 | 286 | 4.28 | 4.0% |
| 4,001–8,000 | 4,000 | 593 | 3.91 | 3.4% |
Word shape
Character lengths across the Swedish top 1,000 — average 5.0.
Characters per word
Phrases that behave like words
1,444 multi-word units earned a rank of their own.
- #76 jag vet I know
- #111 kom igen come on
- #119 vet inte don't know / I don't know
- #135 det finns there is
- #143 jag tror I think
- #209 tror att thinks that / I think that
- #215 tror du do you think
- #218 jag älskar I love
What stands out in Swedish
Read straight off this language's own numbers.
- Just 61 entries cover half of everything said in the Swedish subtitle corpus.
- The ten most common Swedish entries alone account for 23.3% of running text.
- 1,444 of the top 9,901 entries are multi-word phrases that behave like single units — the most common is “jag vet”.
- The longest single word inside the Swedish top 1,000 is “tillräckligt” (12 characters, rank 682).
- Most of the Swedish top 1,000 is short: 5-character words are the single largest group, and the average is 5.0 characters.
- Frequency falls away fast: the whole 4,001–8,000 band together covers only 3.4% of the corpus.
How this Swedish ranking was calculated
Built once, offline, and shipped as static data.
The source is the Swedish side of OpenSubtitles v2024 via OPUS — 14,067,135 tokens of transcribed speech, which is closer to how people talk than a newspaper or Wikipedia corpus would be.
Words are counted using Swedish's own rules rather than by splitting on spaces, so the counts hold up for writing systems that do not put spaces between words. Timings, formatting and speaker names are stripped out first. 230,212 distinct words appear in all; the 9,901 most common are kept.
Repeated word sequences that hold together are promoted into the same ranking as single words, which is why 1,444 entries here are phrases. Each entry carries a raw count, a per-million rate, a Zipf score, and the cumulative share of running text it and everything above it cover.
English glosses come from a translation cascade with per-entry provenance, and anything the translators only echoed back is marked unresolved rather than presented as a translation. 303 of the 9,901 Swedish entries are still unresolved and are never used as game questions.
Common questions about Swedish word frequency
Answered from this language's own corpus.
How many Swedish words do you need to know?
Around 61 Swedish words and phrases cover half of everything said in ordinary speech, and about 1,484 cover 80%. Coverage climbs steeply at first and then flattens: the whole top 9,901 reaches 89.3%, so the last few thousand entries add far less than the first few hundred.
What are the most common Swedish words?
The ten most common Swedish entries are jag (I), det (it / that), är (is), du (you), inte (not), att (that), en (a / one), och (and), vi (we) and har (has). Together they account for 23.3% of all running Swedish text in the corpus.
What is the most common word in Swedish?
The most common Swedish word is “jag”, meaning “I”. It appears 550,087 times across the corpus, which is 3.9% of everything said.
How is this Swedish frequency list calculated?
The list is built from the Swedish side of the OpenSubtitles corpus — 14,067,135 tokens of transcribed dialogue, which reflects spoken language far more closely than a news or encyclopedia corpus. Text is tokenised with Unicode word boundaries, subtitle timing and formatting are stripped, and the 230,212 distinct forms found are ranked by raw count. The top 9,901 are kept, each with a per-million rate, a Zipf score, and its cumulative share of running text.
Does the Swedish list include phrases as well as words?
Yes. 1,444 of the 9,901 Swedish entries are multi-word phrases that repeat tightly enough to behave like single vocabulary items — the most common is “jag vet” (I know), at rank 76. They are ranked alongside single words rather than in a separate list, because knowing them as units is what fluency looks like.
How much Swedish do the top 100 words cover?
The 100 most common Swedish entries cover 56.4% of running text by themselves. That is why they are worth learning first: no other 100 items in the language come close.
Is the Swedish frequency dictionary free to download?
Yes. All 9,598 translated Swedish entries can be exported as JSON, CSV, TSV, Markdown, or an Anki deck straight from the dictionary page, with no account and no sign-up. The data is prebuilt and served as static files.
Corpora of a similar size
Languages whose subtitle corpus is closest to Swedish's.