Floa Words: vocabulary data sources
Research date: 2026-09-20. Recommendation based on provider documentation; no bulk dataset has been imported or pair-level quality measured. Sources can change, so pin the selected release and its licensing metadata during implementation.
Selected source
Decision (2026-09-20): Kaikki/Wiktionary is the selected primary data source for words.floa.stream. The user has approved this source choice. Import raw Wiktextract JSONL snapshots from Kaikki, preserving Wiktionary provenance. Implementation and the first import are pending.
Tatoeba and wordfreq remain optional enrichment candidates, not selected integrations or requirements for the initial feed. Build a reviewed Floa layer for learner categories, sense-specific translations and proficiency levels. No single source reviewed supplies all of those fields consistently across languages.
Implementation contract:
- Stable provider code:
kaikki_wiktionary; retain Wiktionary edition and dump version separately. - Discovery endpoint:
https://kaikki.org/dictionary/rawdata.html; pin each imported snapshot URL and checksum. - Import lemmas, distinct senses, available direct translations, grammatical information and source topic evidence.
- Preserve per-component source links, attribution and license metadata; publish only validated content.
- Assign reviewed levels/categories separately; missing translations or uncertain senses remain unpublished for that language pair.
- Preserve canonical Floa sense IDs and learner progress across source updates.
- Select initial language pairs before the first production import; multilingual availability alone does not establish pair coverage.
Keep Lexicala as the commercial alternative if editorial consistency and reduced curation work justify a paid agreement. FreeDict can supplement specific translation pairs. CEFRLex is relevant to difficulty labeling, but the English and French downloads reviewed carry noncommercial terms and should not enter a commercial catalog without appropriate permission.
Source comparison
| Source | What it provides | Access and coverage | Floa assessment |
|---|---|---|---|
| Kaikki / Wiktextract | Word senses, definitions, translations when available, part of speech, forms, pronunciation links and usage information | Bulk JSONL extracted from Wiktionary editions; broad multilingual coverage | Selected primary source. Extraction and translation coverage vary; normalization and editorial selection remain necessary. |
| Tatoeba | Sentences, translation links, contributor metadata and optional audio metadata | Bulk files and custom sentence-pair exports | Best complementary example source. Match the intended meaning; a sentence containing a word is not automatically a suitable example for its card. |
| wordfreq | Usage-frequency estimates in over 40 languages | Python package including frequency data | Useful ranking input, not a dictionary or proficiency syllabus. |
| FreeDict | Bilingual dictionaries | Downloadable dictionaries and a metadata/download catalog | Evaluate pair by pair; do not assume equal coverage or independent provenance. |
| CEFRLex | Lexical difficulty resources for several European languages | Research resources with language-specific coverage and terms | Potential licensed level enrichment; not the default production source. |
| Lexicala | Human-curated lexical entries, translations, examples, grammatical information and domain labels | JSON API across resources covering 50 languages | Commercial option; obtain terms covering our actual feed display, storage and bulk-import needs. |
Primary dictionary: Kaikki / Wiktionary
The Kaikki index lists English, Polish, German, Spanish, French, Italian, Portuguese, Ukrainian, Japanese, Chinese and many other languages. The English Wiktionary edition describes words from many languages using English glosses; a gloss in English is not automatically a concise translation into the learner’s selected language. Non-English editions, including Polish, are available, with varying extraction completeness.
Use the raw download page. Kaikki explicitly marks the older postprocessed per-language downloads as deprecated. The raw English-edition file is a multi-gigabyte dataset, so process it outside the interactive feed request and emit bounded, validated import batches. Preserve dump date, extractor/schema version, checksum, source page and original source records. The extractor documentation describes JSONL and available fields.
Treat each approved sense as a learning item. Collapse ordinary inflected forms into their lemma unless a form is deliberately taught. Exclude obscure, archaic and unsuitable senses from starter decks. Map source topics to a small reviewed taxonomy such as food, travel, home and work; source categories are evidence rather than ready-made curriculum labels.
English Wiktionary’s copyright page describes CC BY-SA 4.0/GFDL text licensing and exceptions for externally sourced material. The extractor’s MIT software license does not license the dictionary data. Preserve attribution and the applicable content terms; handle audio and quoted examples separately. Prefer Tatoeba or original reviewed examples over assuming every quotation in a dictionary dump is freely reusable.
Examples: Tatoeba
Tatoeba downloads include sentence/translation exports, attribution information, contributor skill data and experimental sentence reviews. Text is generally CC BY 2.0 FR, with a CC0 subset; audio has contributor-selected licenses and must be checked separately. Do not reuse audio with an empty license field.
Proposed selection: direct translation link for the chosen pair, suitable sentence length, natural wording, correct sense, and available attribution. Native-speaker and review metadata can help prioritize editorial review, but do not guarantee correctness. Keep sentence IDs and each side’s provenance. Contributor language proficiency is not the CEFR level of a sentence.
Frequency: wordfreq
The project documentation includes Polish, English, German, French, Spanish, Italian and other languages. The software uses Apache licensing; included data is distributed under CC BY-SA 4.0 with additional attribution notices. Preserve the package’s normalization behavior and credits rather than extracting an uncredited flat word list; the maintainer explicitly discourages standalone CSV redistribution.
The maintainer’s sunset note states that its data reflects usage through about 2021 and will not be updated. It can still help rank ordinary vocabulary, but should not drive claims about current slang or newly common words. Frequency describes word forms, not necessarily individual senses, and must not be converted mechanically into A1–C2 labels. Proposed integration: a pinned Python preprocessing step with frequency provenance retained on imported ranking metadata, subject to checking the packaged notices for the selected use.
Proficiency and categories
The EFLLex download covers English A1–C1, while FLELex offers French resources through C2. Both inspected pages state CC BY-NC-SA 4.0. Do not assume that a downloadable academic dataset is cleared for a commercial Floa product, or that every language has equivalent coverage.
For the first release, commission or perform qualified editorial review of levels and categories for a bounded starter catalog. Record level scheme, level, evidence, reviewer and confidence separately from frequency. Estimated levels must be labeled as estimates; unknown levels stay out of strict level-filtered discovery. Later, license suitable curriculum data where available. The same spelling can have different difficulty depending on sense.
Alternatives
FreeDict documentation places each dictionary’s license in its TEI header. Its download catalog exposes versions, sizes, download links and checksums. Inspect the actual pair before adopting it; some dictionaries derive from other sources and are not independent confirmation of a translation.
Lexicala advertises human-curated lexical resources and EdTech uses. Its published terms also restrict database creation and certain caching and dictionary-display uses without written agreement. Request a license explicitly covering stored vocabulary, translations, examples, card display, adaptation and any offline/export feature. No quote or permission has been obtained; do not equate an API subscription with bulk redistribution rights. Its domain labels are promising category inputs, but verified CEFR coverage was not established by this research.
Integration and initial validation
Proposed pipeline:
Kaikki raw snapshot → normalize lemmas/senses → verify direct translation pairs → review levels/categories → publish to D1
Example and frequency enrichment can be added later through separately selected sources.
Keep bulk parsing in a preprocessing job; use Floa’s existing durable import jobs for bounded publication. Retain attribution at field or source-component level because a definition, translation, example and audio clip may have different origins. Present a compact Sources link on each card.
Do not create German-to-Polish translations merely by joining both words to the same English string: homonyms and sense differences make that unreliable. Prefer documented direct sense-linked translations; send missing or ambiguous pairs to editorial review. Translation models may propose gap fills later, but those are generated candidates rather than source-verified dictionary facts.
For the pilot, choose one or two explicit language pairs and aim for 1,000–2,000 reviewed common senses per pair. This is a proposed release target, not a measured source yield. Assess 100 candidate cards per pair before committing to the importer: translation correctness, sense alignment, example naturalness, rights metadata and coverage by desired level/category. Potential candidates include English–Polish and German–Polish if Polish is the intended learner language; that preference has not been confirmed.
Only enable a language pair once its published inventory has sufficient coverage. “Dictionary contains this language” must not mean “every translation pair is supported.” Keep canonical Floa sense IDs stable through reimports, with source mappings and editorial merge history so learned words never reappear merely because the dataset version changed.