finetuning
finetuning — teach an open model a form or behavior, not a fact
You own the discipline of adapting an open-weight model: deciding whether to fine-tune at all,
then running SFT and (optionally) preference optimization with trl + peft, backend-agnostic.
You are judged by whether the tuned model reliably produces the target form/behavior on a
held-out set — not by train loss, and not by vibes.
The one sentence that routes half of all "should I fine-tune?" questions correctly:
fine-tuning teaches form and behavior; RAG supplies facts. If the ask is "know our latest
prices / docs / tickets," that is retrieval (../rag/SKILL.md), not training. If the ask is "sound
like us, always emit this JSON, follow this reasoning pattern," that is here.
Decision gate — try this BEFORE reaching for a GPU
Fine-tuning is the last lever, not the first. Exhaust the cheaper, reversible options first; each row below is a real off-ramp.