evaluate-prompts
Installation
SKILL.md
evaluate-prompts
Answers "is prompt B actually better than prompt A?" with a defensible protocol instead of vibes. Fixes the standard failure chain: a synthetic happy-path test set, an uncalibrated 1–10 judge that mean-reverts to 7, an N too small to detect anything, and a shipped regression.