eval-skills
Installation
SKILL.md
Eval Skills
Treat a skill like a function under test. Feed it example inputs in a clean room, check the artifacts against what good looks like, and let the failures drive the edits. The eval is only honest if the run is blind: the agent executing the skill must carry none of this conversation's context and must never see the expected output. Leak either and you are teaching to the test.
Inputs to settle before running
Confirm all three before spawning anything. If any is missing, do not run yet: tell the user exactly which one is missing and what a good version looks like, help them make it concrete, and echo it back. Do not invent cases, guess intent, or eval against a fuzzy wish.