magic-linguistic-scope
Installation
SKILL.md
When to Use
- User names a target language for any NLP/LLM project.
- Workflow needs ISO 639-3 + Glottolog identity before touching data.
- Resource-class assessment (Joshi 0-5) is required to pick a strategy.
- Choosing a transfer source (which related language to bootstrap from).
- Determining vitality status (does this need community engagement?).
When NOT to use: the language is already disambiguated, classified, and typology-vector is already in workspace_state.md — proceed to the next specialist.
Thinking Framework — before any data, model, or eval decision
Ask yourself, in order:
- What is this language, exactly? Not "Chinese" — Mandarin (cmn), Cantonese (yue), Wu (wuu), or one of ~20 others. Not "Arabic" — MSA (arb), Egyptian (arz), Levantine (apc), Maghrebi (ary). Macrolanguages silently destroy weeks of work when conflated.
- What resource class is it? Joshi 0-5 changes EVERY downstream choice — tokenizer, eval suite, transfer source, ethics depth.
- What's its typological profile? Word order, morphology type, agreement, tone. These predict what models will get wrong before you measure it.
- What's the best transfer source? Not always English. Often a typologically-closer high-resource language gives 2-5× the transfer gain.
- Is the community engaged? Vitality status (UNESCO/EGIDS) gates how much community involvement is required.