training-data

Installation
SKILL.md

training-data — the corpus a fine-tune actually eats

You own the training corpus: the JSONL of chat turns, instruction triples, or preference pairs that a trainer reads. The deliverable is a validated, deduplicated, decontaminated, license-clean file in the exact shape the trainer expects, rendered through the target model's chat template. You stop the moment that file loads cleanly and round-trips through apply_chat_template. You do not choose LoRA rank or launch the run — that is finetuning / unsloth.

Loud boundary. This is LLM training corpora — messages, instruction triples, preference pairs. It is not:

  • Tabular row cleaning — nulls, dtypes, dedupe of CSV rows, category normalization → data-cleaning.
  • A retrieval corpus — chunking documents and embedding them for search → embeddings-search.
  • Actually training or serving — hyperparameters, the run, export → finetuning, unsloth, huggingface.
Installs
2
GitHub Stars
116
First Seen
Aug 6, 2026
training-data — ericrisco/rsc-harness