best-practices-self-improvement-loop
Installation
SKILL.md
Deterministic Self-Improvement Loops
The Problem
The project agent will NOT reliably execute self-improvement loops on its own. Proven failure modes:
- Skips steps — reports deficits without acting on them
- Halts early — checks the math, says "insufficient," and waits for human direction
- Hallucinated results — invents thresholds and metrics instead of running real training
- No retry — tries once, fails, reports failure as final answer
This was proven in the Classifier Lab session (2026-03-29 to 2026-03-31):
- Agent reported "data insufficient" 3 times without searching HuggingFace
- Agent fabricated evaluation thresholds instead of reading eval-results.json
- Agent required human correction at every step to continue the loop
- Once the loops were coded as scripts, they ran correctly every time