# Third-party data notices This package bundles (and, unlike the research package it was ported from, actually DISTRIBUTES to end-user browsers) three derived data artifacts. This file records what this release actually ships and the attribution/share-alike obligations that travel with it. The full research provenance (fetch dates, pinned revisions, hashes, exact row-count accounting) lives in the `phonetic-foundation` research package's own `NOTICE.md`, not duplicated here in full — this file states only what a DISTRIBUTED-PRODUCT notice needs to say. ## `artifacts/pronunciation-lexicon/wikipron-import.json` Derived from **WikiPron** (Wiktionary pronunciation data), licensed CC BY-SA 4.0 / GFDL 1.1+ (Wiktionary's own dual terms). Commercial reuse is permitted; attribution and a share-alike obligation on the derived data apply. ## `artifacts/pronunciation-lexicon/kaikki-arabic-import.json` Derived from **Kaikki/Wiktextract** (`kaikki.org`, Tatu Ylonen), itself Wiktionary-sourced content under the SAME CC BY-SA 4.0 / GFDL 1.1+ terms. Extraction code is Apache-2.0 (not itself distributed here — only its published data output was used). Citation: Tatu Ylonen, "Wiktextract: Wiktionary as Machine-Readable Structured Data," LREC 2022. ## `artifacts/word-frequency/wordfreq-ar-large.json` Derived from **wordfreq** (Robyn Speer), data licensed CC BY-SA 4.0. Built from Wikipedia, OPUS OpenSubtitles 2018, NewsCrawl 2014 + GlobalVoices, OSCAR web crawl, and pre-2023 Twitter — the Arabic- specific source mixture; Google Books Ngrams and Reddit (which carry their own separate attribution asks) are NOT part of the Arabic data and are not implicated here. Extraction code is Apache-2.0. Data frozen at a 2021 snapshot per the upstream project's own stated sunset. ## `artifacts/learned-spelling-model/variant-source-balanced.json` Originally trained by this project on the two artifacts above plus Claude-authored synthetic sentence pairs — no third-party license obligation beyond the two entries above (the training data those learned costs were derived from), since only the resulting small cost table (numeric weights, no verbatim source text) is distributed here. ## Public attribution surface CC BY-SA's own attribution requirement is satisfied by reasonably accessible credit near the distributed content. Both existing `/legal-notice` (ArabicKeyboard.ai) and `/mentions-legales` (ClavierArabe.ai) pages now carry a short, scoped paragraph — inside each page's existing "Open-source components" section, no new route or redesign — naming all three sources above, their license, and linking back to this exact file (served verbatim at each app's own `/phonetic-conversion/NOTICE.txt`, the same established convention this project already uses for `/spellcheck/NOTICE.txt` and `/fonts/text-card/NOTICE.txt`). This file is the one source of truth; keep the two copies under `apps/*/public/phonetic-conversion/NOTICE.txt` byte-identical to it if it changes.