gpt-lab
Installation
SKILL.md
STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
GPT Lab
Benchmark and compare small GPTs trained by /create-gpt against prompted alternatives.
Answers the key question: "Is fine-tuning worth it for this task?"
Minimum Data for Meaningful Benchmarks
Benchmarking a model trained on < 1,000 examples will produce misleading results. The model hasn't learned the task — you're benchmarking noise.
Evidence: sparta_stress_grading (246 training samples) → 33.7% shadow agreement.
No amount of benchmarking fixes an undertrained model. Collect more labels first.
See /create-gpt SKILL.md "Minimum Training Data" section for thresholds.