NoaLingua for Chrome
Add to Chrome — Free Add to Chrome

Original research

The NoaLingua Language Data Index

What the open language-data ecosystem actually covers, measured rather than assumed

Only 7 of the 28 languages NoaLingua supports are reached by all 5 of the open data sources it draws on. This page publishes the full measurement — availability, pack size in megabytes, and paradigm row counts — as a citable dataset, because to ship the product we had to measure something nobody had published.

Why this exists

NoaLingua works offline, which means every dictionary, grammar table and example sentence it shows has to be downloaded from an open corpus first. Before we could promise that for a language, we had to find out whether the data existed — so we probed each corpus, language by language, and measured what came back.

That measurement turned out to be the interesting part. UniMorph documents UniMorph. kaikki documents kaikki. Tatoeba documents Tatoeba. None of them measures the intersection, and none of them states what any of it weighs on disk. We had to assemble that to ship, so we are publishing it: 140 measured cells, 28 languages against 5 sources, free to reuse.

Six findings

1. Coverage is far from universal — and only one source is

Of the five sources, exactly one reaches every language: Tatoeba, at 28 of 28. Human-written example sentences turn out to be the most evenly distributed open language resource there is, and the frequency lists the least — 7 languages, 25 %.

Each of the five open sources, and how many of the 28 languages it reaches
Source Corpus Languages Share Licence
Grammar tables UniMorph 24 of 28 86 % CC BY-SA 3.0
Dictionary and IPA kaikki.org (Wiktionary extraction) 24 of 28 86 % CC BY-SA 4.0
Example sentences Tatoeba 28 of 28 100 % CC BY 2.0 FR
Word frequency Hermit Dave / OpenSubtitles 2018 7 of 28 25 % MIT
Punctuation restoration anchor-flux/punct-pcs-47lang 22 of 28 79 % Apache 2.0

2. 7 of 28 languages have complete coverage

Only English, Spanish, French, German, Italian, Portuguese, Russian are reached by all five — 25 % of the set. Across the whole grid, 105 of 140 cells are filled. The distribution is the shape of the problem:

  • 5 of 5 sources — 7 languages
  • 4 of 5 sources — 11 languages
  • 3 of 5 sources — 8 languages
  • 1 of 5 sources — 2 languages

The two languages at the bottom are Vietnamese and Thai, reached by example sentences and nothing else. Both are analytic languages, and the two absences compound: there is nothing to tabulate morphologically, and no Wiktionary extract exists. That is the thinnest offline coverage in the set, and it is not a coincidence of neglect — it is what happens when a language's information lives somewhere the tooling is not looking.

3. Most absences are honest, and one is not

Four languages have no grammar dataset. Three of those absences — Vietnamese, Thai, Chinese — are linguistic facts rather than gaps: these are analytic or isolating languages, and there is genuinely nothing to conjugate or decline. Recording them as "missing" would misrepresent the language, not the corpus.

The fourth is different. Finnish has a dataset; upstream splits it across two files where every other language uses one, and a downloader that assumes one file silently fetches half of it. For a language this heavily inflected, that is the single most consequential gap in the set — and it is a tooling limitation wearing the costume of a data limitation, which is exactly the kind of thing a coverage table hides unless it is asked to distinguish them.

4. Dictionary weight is extraordinarily concentrated

Across the 19 languages whose dictionary packs we measured, the total is 1,416 MB. English alone accounts for 34 % of it, at 475 MB. The five largest together — English, Chinese, German, Spanish, Russian — account for 62 %.

Grammar data is a different order of magnitude entirely: 349 MB across 24 languages, against 1,416 MB of dictionary. Open language data is, overwhelmingly, lexicon.

5. The dictionary-to-grammar ratio spans forty-five fold

For each language where both packs were measured, dividing dictionary megabytes by grammar megabytes gives a rough answer to a real linguistic question: where does this language keep its information? The range is remarkable. Japanese sits at 45.0× — almost everything in the lexicon. Turkish sits at 1.0× — an almost exactly even split, which is what agglutination looks like measured in bytes: one stem generating a very large number of forms, each of which has to be stored.

This is the finding we did not expect and cannot fully explain. It is offered as a measurement, not a theory. Pack size is a proxy for information content and a crude one — compression, extraction method and corpus completeness all confound it — so the ratio is a place to start looking, not a conclusion.

6. Where paradigms were counted, they are enormous

For 6 languages we counted the rows rather than trusting the file size: 3,433,493 inflected forms across five distinct datasets, since Serbian and Croatian share one. Two of those datasets — French and Italian — contain verbs and nothing else. Not a single noun or adjective, despite both being languages with substantial nominal morphology.

A coverage table that reports those two as "grammar: yes" is technically correct and practically misleading, which is why this dataset carries the part-of-speech spread as its own column rather than folding it into a tick.

The data

One row per language. ● present, ○ absent. Sizes are measured megabytes; a blank means the pack exists but has never been measured, which is reported as unknown rather than guessed at.

The Language Data Index: 28 languages against 5 open data sources
Language Grammar tablesDictionary and IPAExample sentencesWord frequencyPunctuation restoration Of 5 Grammar MB Dictionary MB Paradigm rows
English English present present present present present 5 18 475 —
French Français present present present present present 5 13 54 367,732 (verbs only)
German Deutsch present present present present present 5 19 91 519,143
Italian Italiano present present present present present 5 20 71 509,574 (verbs only)
Portuguese Português present present present present present 5 11 51 —
Russian Русский present present present present present 5 25 85 —
Spanish Español present present present present present 5 50 85 1,196,245
Arabic العربية present present present absent present 4 7 47 —
Dutch Nederlands present present present absent present 4 2 28 —
Greek Ελληνικά present present present absent present 4 11 23 —
Hindi हिन्दी present present present absent present 4 4 16 —
Hungarian Magyar present present present absent present 4 37 — —
Japanese 日本語 present present present absent present 4 1 45 —
Korean 한국어 present present present absent present 4 12 19 —
Polish Polski present present present absent present 4 7 72 —
Romanian Română present present present absent present 4 3 27 —
Turkish Türkçe present present present absent present 4 32 32 —
Ukrainian Українська present present present absent present 4 1 — —
Chinese 中文 absent present present absent present 3 — 145 —
Croatian Hrvatski present absent present absent present 3 31 — 840,799
Czech Čeština present present present absent absent 3 5 19 —
Danish Dansk present present present absent absent 3 1 — —
Finnish Suomi absent present present absent present 3 — — —
Norwegian Norsk present present present absent absent 3 3 — —
Serbian Српски present absent present absent present 3 31 — 840,799
Swedish Svenska present present present absent absent 3 5 31 —
Thai ไทย absent absent present absent absent 1 — — —
Vietnamese Tiếng Việt absent absent present absent absent 1 — — —

Download and reuse

The full dataset, in both formats, under Creative Commons Attribution 4.0. Use it, republish it, build on it — the only condition is attribution.

Cite this

Pejić, N. (2026). NoaLingua Language Data Index (v1.0) [Data set]. NoaSolutions. https://noalingua.com/research/language-data-index/

Download the BibTeX citation.

Method, and what it does not show

Availability was established by probing each corpus's own release URL per language and recording the HTTP status; a 404 is recorded as a documented absence, not as an unknown. Pack sizes come from HEAD requests against the same URLs. Row counts were read from the downloaded files. Everything was measured on 2026-08-17 against the live services, not read off a published specification.

The limits are worth stating plainly, because they are real:

  • This is a snapshot. All five corpora are actively maintained. A language absent on 2026-08-17 may be present now.
  • The language set is not neutral. These are the 28 languages NoaLingua supports, chosen for a product. They over-represent European languages, and the coverage picture for languages outside this set is almost certainly worse than what is shown here, not better.
  • Size is a proxy, not a measure of quality. A large pack can be redundant and a small one dense. The ratio finding above is the most exposed to this and should be treated as the most provisional.
  • 5 dictionaries were never sized. Ukrainian, Hungarian, Finnish, Danish, Norwegian have a pack that exists and a size we have not measured. They are blank above rather than estimated.
  • Two languages share one dataset. UniMorph publishes a single Serbo-Croatian file, so Serbian and Croatian carry identical grammar figures and the paradigm total counts it once. It cannot distinguish the two standards' spelling, which is a limitation of the upstream data and not of the measurement.

What the underlying registries mean for a specific language is on that language's page under supported languages. The licences and terms of each corpus are on data sources. What the product cannot do with any of it is on the limits page.