NoaLingua for Chrome
Add to Chrome — Free Add to Chrome

Research

Research NoaLingua publishes

Things we had to measure to build the product, released as data

Building an offline language-learning extension meant measuring which open corpora actually cover which languages, and what they weigh. Nobody had published that intersection, so we are releasing it. Everything here is free to reuse under Creative Commons Attribution, with the method and its limits stated in full.

Published datasets

  • The Language Data Index

    Measured coverage of 28 languages across 5 open language-data corpora, with pack sizes and paradigm counts. 7 languages are reached by all five.

    Published 2026-09-09 · measured 2026-08-17 · CC BY 4.0

How to use it

Everything on this page is released under Creative Commons Attribution 4.0. Republish it, chart it, correct it, build something better on top of it — the only condition is that you say where it came from. Each dataset carries its own citation line and ships as both CSV and JSON.

Corrections are genuinely wanted. If a probe was wrong, or a corpus has moved since we measured it, tell us and the dataset gets a new version with the change recorded rather than a quiet edit.