magic-linguistic-historical
When to Use
- Class 0-1 target language; need bilingual lexicon from related-language cognates.
- Bootstrapping bitext from Swadesh-list pairs.
- Validating typological-distance recommendations from
magic-linguistic-scope. - Cognate-based data augmentation between related-language pairs.
When NOT to use: Class 3+ language → standard data ample; cognate bootstrap not needed. Pure typology lookup → magic-linguistic-scope.
Stance
Comparative-historical linguistics is decades-mature. For low-resource ML, the operationalizable primitives are: cognate sets, Swadesh lists, sound correspondences. Use them as cheap data-augmentation tools when you have a related higher-resource language to draw from.
What's worth knowing
-
Cognate sets — etymologically-related word pairs across related languages. Spanish "noche" ↔ Italian "notte" ↔ French "nuit" ↔ Portuguese "noite" — same Latin source, predictable sound shifts. For class 0-1, cognate sets bootstrap bilingual lexicons cheaply.
-
Swadesh lists — 100-word / 200-word core-vocabulary lists. Standard starting point for class 0 bilingual-lexicon work. NOT a complete lexicon; a starting bootstrap.