Original research
The NoaLingua Language Data Index
What the open language-data ecosystem actually covers, measured rather than assumed
Only 7 of the 28 languages NoaLingua supports are reached by all 5 of the open data sources it draws on. This page publishes the full measurement — availability, pack size in megabytes, and paradigm row counts — as a citable dataset, because to ship the product we had to measure something nobody had published.
Why this exists
NoaLingua works offline, which means every dictionary, grammar table and example sentence it shows has to be downloaded from an open corpus first. Before we could promise that for a language, we had to find out whether the data existed — so we probed each corpus, language by language, and measured what came back.
That measurement turned out to be the interesting part. UniMorph documents UniMorph. kaikki documents kaikki. Tatoeba documents Tatoeba. None of them measures the intersection, and none of them states what any of it weighs on disk. We had to assemble that to ship, so we are publishing it: 140 measured cells, 28 languages against 5 sources, free to reuse.
Six findings
1. Coverage is far from universal — and only one source is
Of the five sources, exactly one reaches every language: Tatoeba, at 28 of 28. Human-written example sentences turn out to be the most evenly distributed open language resource there is, and the frequency lists the least — 7 languages, 25 %.
| Source | Corpus | Languages | Share | Licence |
|---|---|---|---|---|
| Grammar tables | UniMorph | 24 of 28 | 86 % | CC BY-SA 3.0 |
| Dictionary and IPA | kaikki.org (Wiktionary extraction) | 24 of 28 | 86 % | CC BY-SA 4.0 |
| Example sentences | Tatoeba | 28 of 28 | 100 % | CC BY 2.0 FR |
| Word frequency | Hermit Dave / OpenSubtitles 2018 | 7 of 28 | 25 % | MIT |
| Punctuation restoration | anchor-flux/punct-pcs-47lang | 22 of 28 | 79 % | Apache 2.0 |
2. 7 of 28 languages have complete coverage
Only English, Spanish, French, German, Italian, Portuguese, Russian are reached by all five — 25 % of the set. Across the whole grid, 105 of 140 cells are filled. The distribution is the shape of the problem:
- 5 of 5 sources — 7 languages
- 4 of 5 sources — 11 languages
- 3 of 5 sources — 8 languages
- 1 of 5 sources — 2 languages
The two languages at the bottom are Vietnamese and Thai, reached by example sentences and nothing else. Both are analytic languages, and the two absences compound: there is nothing to tabulate morphologically, and no Wiktionary extract exists. That is the thinnest offline coverage in the set, and it is not a coincidence of neglect — it is what happens when a language's information lives somewhere the tooling is not looking.
3. Most absences are honest, and one is not
Four languages have no grammar dataset. Three of those absences — Vietnamese, Thai, Chinese — are linguistic facts rather than gaps: these are analytic or isolating languages, and there is genuinely nothing to conjugate or decline. Recording them as "missing" would misrepresent the language, not the corpus.
The fourth is different. Finnish has a dataset; upstream splits it across two files where every other language uses one, and a downloader that assumes one file silently fetches half of it. For a language this heavily inflected, that is the single most consequential gap in the set — and it is a tooling limitation wearing the costume of a data limitation, which is exactly the kind of thing a coverage table hides unless it is asked to distinguish them.
4. Dictionary weight is extraordinarily concentrated
Across the 19 languages whose dictionary packs we measured, the total is 1,416 MB. English alone accounts for 34 % of it, at 475 MB. The five largest together — English, Chinese, German, Spanish, Russian — account for 62 %.
Grammar data is a different order of magnitude entirely: 349 MB across 24 languages, against 1,416 MB of dictionary. Open language data is, overwhelmingly, lexicon.
5. The dictionary-to-grammar ratio spans forty-five fold
For each language where both packs were measured, dividing dictionary megabytes by grammar megabytes gives a rough answer to a real linguistic question: where does this language keep its information? The range is remarkable. Japanese sits at 45.0× — almost everything in the lexicon. Turkish sits at 1.0× — an almost exactly even split, which is what agglutination looks like measured in bytes: one stem generating a very large number of forms, each of which has to be stored.
This is the finding we did not expect and cannot fully explain. It is offered as a measurement, not a theory. Pack size is a proxy for information content and a crude one — compression, extraction method and corpus completeness all confound it — so the ratio is a place to start looking, not a conclusion.
6. Where paradigms were counted, they are enormous
For 6 languages we counted the rows rather than trusting the file size: 3,433,493 inflected forms across five distinct datasets, since Serbian and Croatian share one. Two of those datasets — French and Italian — contain verbs and nothing else. Not a single noun or adjective, despite both being languages with substantial nominal morphology.
A coverage table that reports those two as "grammar: yes" is technically correct and practically misleading, which is why this dataset carries the part-of-speech spread as its own column rather than folding it into a tick.
The data
One row per language. ● present, ○ absent. Sizes are measured megabytes; a blank means the pack exists but has never been measured, which is reported as unknown rather than guessed at.
| Language | Grammar tables | Dictionary and IPA | Example sentences | Word frequency | Punctuation restoration | Of 5 | Grammar MB | Dictionary MB | Paradigm rows |
|---|---|---|---|---|---|---|---|---|---|
| English English | present | present | present | present | present | 5 | 18 | 475 | — |
| French Français | present | present | present | present | present | 5 | 13 | 54 | 367,732 (verbs only) |
| German Deutsch | present | present | present | present | present | 5 | 19 | 91 | 519,143 |
| Italian Italiano | present | present | present | present | present | 5 | 20 | 71 | 509,574 (verbs only) |
| Portuguese Português | present | present | present | present | present | 5 | 11 | 51 | — |
| Russian Русский | present | present | present | present | present | 5 | 25 | 85 | — |
| Spanish Español | present | present | present | present | present | 5 | 50 | 85 | 1,196,245 |
| Arabic العربية | present | present | present | absent | present | 4 | 7 | 47 | — |
| Dutch Nederlands | present | present | present | absent | present | 4 | 2 | 28 | — |
| Greek Ελληνικά | present | present | present | absent | present | 4 | 11 | 23 | — |
| Hindi हिन्दी | present | present | present | absent | present | 4 | 4 | 16 | — |
| Hungarian Magyar | present | present | present | absent | present | 4 | 37 | — | — |
| Japanese 日本語 | present | present | present | absent | present | 4 | 1 | 45 | — |
| Korean 한국어 | present | present | present | absent | present | 4 | 12 | 19 | — |
| Polish Polski | present | present | present | absent | present | 4 | 7 | 72 | — |
| Romanian Română | present | present | present | absent | present | 4 | 3 | 27 | — |
| Turkish Türkçe | present | present | present | absent | present | 4 | 32 | 32 | — |
| Ukrainian Українська | present | present | present | absent | present | 4 | 1 | — | — |
| Chinese 中文 | absent | present | present | absent | present | 3 | — | 145 | — |
| Croatian Hrvatski | present | absent | present | absent | present | 3 | 31 | — | 840,799 |
| Czech Čeština | present | present | present | absent | absent | 3 | 5 | 19 | — |
| Danish Dansk | present | present | present | absent | absent | 3 | 1 | — | — |
| Finnish Suomi | absent | present | present | absent | present | 3 | — | — | — |
| Norwegian Norsk | present | present | present | absent | absent | 3 | 3 | — | — |
| Serbian Српски | present | absent | present | absent | present | 3 | 31 | — | 840,799 |
| Swedish Svenska | present | present | present | absent | absent | 3 | 5 | 31 | — |
| Thai ไทย | absent | absent | present | absent | absent | 1 | — | — | — |
| Vietnamese Tiếng Việt | absent | absent | present | absent | absent | 1 | — | — | — |
Download and reuse
The full dataset, in both formats, under Creative Commons Attribution 4.0. Use it, republish it, build on it — the only condition is attribution.
- language-data-index.csv — one row per language, 28 rows
- language-data-index.json — the same data with the aggregate figures
- language-data-index.bib — a copy-ready BibTeX record
Cite this
Pejić, N. (2026). NoaLingua Language Data Index (v1.0) [Data set]. NoaSolutions. https://noalingua.com/research/language-data-index/
Method, and what it does not show
Availability was established by probing each corpus's own release URL per language and recording the HTTP status; a 404 is recorded as a documented absence, not as an unknown. Pack sizes come from HEAD requests against the same URLs. Row counts were read from the downloaded files. Everything was measured on 2026-08-17 against the live services, not read off a published specification.
The limits are worth stating plainly, because they are real:
- This is a snapshot. All five corpora are actively maintained. A language absent on 2026-08-17 may be present now.
- The language set is not neutral. These are the 28 languages NoaLingua supports, chosen for a product. They over-represent European languages, and the coverage picture for languages outside this set is almost certainly worse than what is shown here, not better.
- Size is a proxy, not a measure of quality. A large pack can be redundant and a small one dense. The ratio finding above is the most exposed to this and should be treated as the most provisional.
- 5 dictionaries were never sized. Ukrainian, Hungarian, Finnish, Danish, Norwegian have a pack that exists and a size we have not measured. They are blank above rather than estimated.
- Two languages share one dataset. UniMorph publishes a single Serbo-Croatian file, so Serbian and Croatian carry identical grammar figures and the paradigm total counts it once. It cannot distinguish the two standards' spelling, which is a limitation of the upstream data and not of the measurement.
What the underlying registries mean for a specific language is on that language's page under supported languages. The licences and terms of each corpus are on data sources. What the product cannot do with any of it is on the limits page.
The product this measurement was taken for
Free, with every feature. No account, nothing to cancel, and your deck stays on your machine.
Add to Chrome — Free Add to Chrome Opens the Chrome Web Store listing NoaLingua — beyond dual subs: video, web, PDF — free, no account, Chrome 138 or newer.- Free, and not a trial Every feature on the free plan — both PDF modes, shadowing, all 30 practice formats.
- No account, no email Nothing to sign up for and nothing to cancel. There is no login screen at all.
- Two dropdowns to set up Pick what you are learning and what you want translations in. That is the whole setup.
- Your words stay yours Export to Anki, CSV, a spreadsheet or a full JSON profile — free on every plan.