---
title: "The Language Data Index: 28 languages · NoaLingua"
description: "Measured coverage of 28 languages across five open language-data sources, with pack sizes and paradigm counts. Free to reuse under CC BY 4.0, as CSV and JSON."
url: "https://noalingua.com/research/language-data-index/"
language: "en"
modified: "2026-09-10"
license: "CC BY 4.0 — https://creativecommons.org/licenses/by/4.0/"
---
Original research

# The NoaLingua Language Data Index

What the open language-data ecosystem actually covers, measured rather than assumed

Only 7 of the 28 languages NoaLingua supports are reached by all 5 of the open data sources it draws on. This page publishes the full measurement — availability, pack size in megabytes, and paradigm row counts — as a citable dataset, because to ship the product we had to measure something nobody had published.

## Why this exists

NoaLingua works offline, which means every dictionary, grammar table and example sentence it shows has to be downloaded from an open corpus first. Before we could promise that for a language, we had to find out whether the data existed — so we probed each corpus, language by language, and measured what came back.

That measurement turned out to be the interesting part. UniMorph documents UniMorph. kaikki documents kaikki. Tatoeba documents Tatoeba. None of them measures the *intersection*, and none of them states what any of it weighs on disk. We had to assemble that to ship, so we are publishing it: 140 measured cells, 28 languages against 5 sources, free to reuse.

## Six findings

### 1. Coverage is far from universal — and only one source is

Of the five sources, exactly one reaches every language: **Tatoeba, at 28 of 28**. Human-written example sentences turn out to be the most evenly distributed open language resource there is, and the frequency lists the least — 7 languages, 25 %.

| Source | Corpus | Languages | Share | Licence |
| --- | --- | --- | --- | --- |
| Grammar tables | [UniMorph](https://unimorph.github.io/) | 24 of 28 | 86 % | CC BY-SA 3.0 |
| Dictionary and IPA | [kaikki.org (Wiktionary extraction)](https://kaikki.org/) | 24 of 28 | 86 % | CC BY-SA 4.0 |
| Example sentences | [Tatoeba](https://tatoeba.org/) | 28 of 28 | 100 % | CC BY 2.0 FR |
| Word frequency | [Hermit Dave / OpenSubtitles 2018](https://github.com/hermitdave/FrequencyWords) | 7 of 28 | 25 % | MIT |
| Punctuation restoration | [anchor-flux/punct-pcs-47lang](https://huggingface.co/) | 22 of 28 | 79 % | Apache 2.0 |

### 2. 7 of 28 languages have complete coverage

Only English, Spanish, French, German, Italian, Portuguese, Russian are reached by all five — 25 % of the set. Across the whole grid, 105 of 140 cells are filled. The distribution is the shape of the problem:

- **5 of 5 sources** — 7 languages
- **4 of 5 sources** — 11 languages
- **3 of 5 sources** — 8 languages
- **1 of 5 sources** — 2 languages

The two languages at the bottom are **Vietnamese and Thai**, reached by example sentences and nothing else. Both are analytic languages, and the two absences compound: there is nothing to tabulate morphologically, *and* no Wiktionary extract exists. That is the thinnest offline coverage in the set, and it is not a coincidence of neglect — it is what happens when a language's information lives somewhere the tooling is not looking.

### 3. Most absences are honest, and one is not

Four languages have no grammar dataset. Three of those absences — Vietnamese, Thai, Chinese — are linguistic facts rather than gaps: these are analytic or isolating languages, and there is genuinely nothing to conjugate or decline. Recording them as "missing" would misrepresent the language, not the corpus.

The fourth is different. Finnish has a dataset; upstream splits it across two files where every other language uses one, and a downloader that assumes one file silently fetches half of it. For a language this heavily inflected, that is the single most consequential gap in the set — and it is a tooling limitation wearing the costume of a data limitation, which is exactly the kind of thing a coverage table hides unless it is asked to distinguish them.

### 4. Dictionary weight is extraordinarily concentrated

Across the 19 languages whose dictionary packs we measured, the total is **1,416 MB**. English alone accounts for **34 %** of it, at 475 MB. The five largest together — English, Chinese, German, Spanish, Russian — account for **62 %**.

Grammar data is a different order of magnitude entirely: 349 MB across 24 languages, against 1,416 MB of dictionary. Open language data is, overwhelmingly, lexicon.

### 5. The dictionary-to-grammar ratio spans forty-five fold

For each language where both packs were measured, dividing dictionary megabytes by grammar megabytes gives a rough answer to a real linguistic question: where does this language keep its information? The range is remarkable. **Japanese** sits at 45.0× — almost everything in the lexicon. **Turkish** sits at 1.0× — an almost exactly even split, which is what agglutination looks like measured in bytes: one stem generating a very large number of forms, each of which has to be stored.

This is the finding we did not expect and cannot fully explain. It is offered as a measurement, not a theory. Pack size is a proxy for information content and a crude one — compression, extraction method and corpus completeness all confound it — so the ratio is a place to start looking, not a conclusion.

### 6. Where paradigms were counted, they are enormous

For 6 languages we counted the rows rather than trusting the file size: **3,433,493 inflected forms** across five distinct datasets, since Serbian and Croatian share one. Two of those datasets — French and Italian — contain *verbs and nothing else*. Not a single noun or adjective, despite both being languages with substantial nominal morphology.

A coverage table that reports those two as "grammar: yes" is technically correct and practically misleading, which is why this dataset carries the part-of-speech spread as its own column rather than folding it into a tick.

## The data

One row per language. ● present, ○ absent. Sizes are measured megabytes; a blank means the pack exists but has never been measured, which is reported as unknown rather than guessed at.

| Language | Grammar tables | Dictionary and IPA | Example sentences | Word frequency | Punctuation restoration | Of 5 | Grammar MB | Dictionary MB | Paradigm rows |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| English English | ● | ● | ● | ● | ● | 5 | 18 | 475 | — |
| French Français | ● | ● | ● | ● | ● | 5 | 13 | 54 | 367,732 (verbs only) |
| German Deutsch | ● | ● | ● | ● | ● | 5 | 19 | 91 | 519,143 |
| Italian Italiano | ● | ● | ● | ● | ● | 5 | 20 | 71 | 509,574 (verbs only) |
| Portuguese Português | ● | ● | ● | ● | ● | 5 | 11 | 51 | — |
| Russian Русский | ● | ● | ● | ● | ● | 5 | 25 | 85 | — |
| Spanish Español | ● | ● | ● | ● | ● | 5 | 50 | 85 | 1,196,245 |
| Arabic العربية | ● | ● | ● | ○ | ● | 4 | 7 | 47 | — |
| Dutch Nederlands | ● | ● | ● | ○ | ● | 4 | 2 | 28 | — |
| Greek Ελληνικά | ● | ● | ● | ○ | ● | 4 | 11 | 23 | — |
| Hindi हिन्दी | ● | ● | ● | ○ | ● | 4 | 4 | 16 | — |
| Hungarian Magyar | ● | ● | ● | ○ | ● | 4 | 37 | — | — |
| Japanese 日本語 | ● | ● | ● | ○ | ● | 4 | 1 | 45 | — |
| Korean 한국어 | ● | ● | ● | ○ | ● | 4 | 12 | 19 | — |
| Polish Polski | ● | ● | ● | ○ | ● | 4 | 7 | 72 | — |
| Romanian Română | ● | ● | ● | ○ | ● | 4 | 3 | 27 | — |
| Turkish Türkçe | ● | ● | ● | ○ | ● | 4 | 32 | 32 | — |
| Ukrainian Українська | ● | ● | ● | ○ | ● | 4 | 1 | — | — |
| Chinese 中文 | ○ | ● | ● | ○ | ● | 3 | — | 145 | — |
| Croatian Hrvatski | ● | ○ | ● | ○ | ● | 3 | 31 | — | 840,799 |
| Czech Čeština | ● | ● | ● | ○ | ○ | 3 | 5 | 19 | — |
| Danish Dansk | ● | ● | ● | ○ | ○ | 3 | 1 | — | — |
| Finnish Suomi | ○ | ● | ● | ○ | ● | 3 | — | — | — |
| Norwegian Norsk | ● | ● | ● | ○ | ○ | 3 | 3 | — | — |
| Serbian Српски | ● | ○ | ● | ○ | ● | 3 | 31 | — | 840,799 |
| Swedish Svenska | ● | ● | ● | ○ | ○ | 3 | 5 | 31 | — |
| Thai ไทย | ○ | ○ | ● | ○ | ○ | 1 | — | — | — |
| Vietnamese Tiếng Việt | ○ | ○ | ● | ○ | ○ | 1 | — | — | — |

## Download and reuse

The full dataset, in both formats, under [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/). Use it, republish it, build on it — the only condition is attribution.

- [language-data-index.csv](https://noalingua.com/research/language-data-index.csv) — one row per language, 28 rows
- [language-data-index.json](https://noalingua.com/research/language-data-index.json) — the same data with the aggregate figures
- [language-data-index.bib](https://noalingua.com/research/language-data-index.bib) — a copy-ready BibTeX record

## Cite this

Pejić, N. (2026). NoaLingua Language Data Index (v1.0) [Data set]. NoaSolutions. https://noalingua.com/research/language-data-index/

[Download the BibTeX citation](https://noalingua.com/research/language-data-index.bib).

## Method, and what it does not show

Availability was established by probing each corpus's own release URL per language and recording the HTTP status; a 404 is recorded as a documented absence, not as an unknown. Pack sizes come from HEAD requests against the same URLs. Row counts were read from the downloaded files. Everything was measured on 2026-08-17 against the live services, not read off a published specification.

The limits are worth stating plainly, because they are real:

- **This is a snapshot.** All five corpora are actively maintained. A language absent on 2026-08-17 may be present now.
- **The language set is not neutral.** These are the 28 languages NoaLingua supports, chosen for a product. They over-represent European languages, and the coverage picture for languages outside this set is almost certainly worse than what is shown here, not better.
- **Size is a proxy, not a measure of quality.** A large pack can be redundant and a small one dense. The ratio finding above is the most exposed to this and should be treated as the most provisional.
- **5 dictionaries were never sized.** Ukrainian, Hungarian, Finnish, Danish, Norwegian have a pack that exists and a size we have not measured. They are blank above rather than estimated.
- **Two languages share one dataset.** UniMorph publishes a single Serbo-Croatian file, so Serbian and Croatian carry identical grammar figures and the paradigm total counts it once. It cannot distinguish the two standards' spelling, which is a limitation of the upstream data and not of the measurement.

What the underlying registries mean for a specific language is on that language's page under [supported languages](https://noalingua.com/supported-languages/). The licences and terms of each corpus are on [data sources](https://noalingua.com/data-sources/). What the product cannot do with any of it is on [the limits page](https://noalingua.com/limits/).

## The product this measurement was taken for

Free, with every feature. No account, nothing to cancel, and your deck stays on your machine.

[Add to Chrome — Free Add to Chrome](https://chromewebstore.google.com/detail/ocimolffhgpeeoifjpjhgikgfbnmlkdh) Opens the Chrome Web Store listing **NoaLingua — beyond dual subs: video, web, PDF** — free, no account, Chrome 138 or newer.

- [Free, and not a trial Every feature on the free plan — both PDF modes, shadowing, all 30 practice formats.](https://noalingua.com/pricing/)
- [No account, no email Nothing to sign up for and nothing to cancel. There is no login screen at all.](https://noalingua.com/privacy/)
- [Two dropdowns to set up Pick what you are learning and what you want translations in. That is the whole setup.](https://noalingua.com/getting-started/)
- [Your words stay yours Export to Anki, CSV, a spreadsheet or a full JSON profile — free on every plan.](https://noalingua.com/docs/export-import/)

[What happens after you install](https://noalingua.com/getting-started/) · [Pricing](https://noalingua.com/pricing/) · [What it cannot do](https://noalingua.com/limits/)
