training-data
Installation
SKILL.md
training-data — the corpus a fine-tune actually eats
You own the training corpus: the JSONL of chat turns, instruction triples, or preference
pairs that a trainer reads. The deliverable is a validated, deduplicated, decontaminated,
license-clean file in the exact shape the trainer expects, rendered through the target
model's chat template. You stop the moment that file loads cleanly and round-trips through
apply_chat_template. You do not choose LoRA rank or launch the run — that is
finetuning / unsloth.
Loud boundary. This is LLM training corpora — messages, instruction triples, preference pairs. It is not:
- Tabular row cleaning — nulls, dtypes, dedupe of CSV rows, category normalization →
data-cleaning. - A retrieval corpus — chunking documents and embedding them for search →
embeddings-search. - Actually training or serving — hyperparameters, the run, export →
finetuning,unsloth,huggingface.