Research
Research NoaLingua publishes
Things we had to measure to build the product, released as data
Building an offline language-learning extension meant measuring which open corpora actually cover which languages, and what they weigh. Nobody had published that intersection, so we are releasing it. Everything here is free to reuse under Creative Commons Attribution, with the method and its limits stated in full.
Published datasets
-
The Language Data Index
Measured coverage of 28 languages across 5 open language-data corpora, with pack sizes and paradigm counts. 7 languages are reached by all five.
How to use it
Everything on this page is released under Creative Commons Attribution 4.0. Republish it, chart it, correct it, build something better on top of it — the only condition is that you say where it came from. Each dataset carries its own citation line and ships as both CSV and JSON.
Corrections are genuinely wanted. If a probe was wrong, or a corpus has moved since we measured it, tell us and the dataset gets a new version with the change recorded rather than a quiet edit.
Try NoaLingua on the next video you were going to watch anyway.
Free, with every feature. No account, nothing to cancel, and your deck stays on your machine.
Add to Chrome — Free Add to Chrome Opens the Chrome Web Store listing NoaLingua — beyond dual subs: video, web, PDF — free, no account, Chrome 138 or newer.- Free, and not a trial Every feature on the free plan — both PDF modes, shadowing, all 30 practice formats.
- No account, no email Nothing to sign up for and nothing to cancel. There is no login screen at all.
- Two dropdowns to set up Pick what you are learning and what you want translations in. That is the whole setup.
- Your words stay yours Export to Anki, CSV, a spreadsheet or a full JSON profile — free on every plan.