nlp
Installation
SKILL.md
nlp — tokenize the text, pick the model type, pick the metric
You own the language-modeling discipline: how raw text becomes tokens, which transformer architecture fits a task, and which metric actually tells you whether it worked. When the question is "which tokenizer," "BERT or GPT or T5 for this," "why does my Catalan text cost 3× the tokens," or "is this BLEU score meaningful," this is the skill. You stop at retrieval, the RAG loop, prompt wording, and the training step itself — those route out (below).
Route out first (loud — do not duplicate these)
- Retrieval embeddings + vector search (which embedding model, chunking, recall@k, rerank)
→
../embeddings-search/SKILL.md. Sentence embeddings live here as a task; using them to retrieve is theirs. - The retrieve → prompt → generate → answer loop and groundedness →
../rag/SKILL.md. - Prompt wording / few-shot / system prompts →
../prompt-engineering/SKILL.md. - Training the network (LoRA/SFT, trainer loop, PyTorch) →
../finetuning/SKILL.md - The training corpus itself (JSONL messages, label sets) →
../training-data/SKILL.md.