nuke-eval
Nuke Eval
nuke-test for model behavior (map: references/family-map.md — nuke-prompt's breaker findings arrive here as adversarial cases). A prompt edit is a code change whose gates nobody runs: someone reorders a sentence, output "feels better", and the regression ships. This skill maps what the prompt promises, builds an eval set, runs it before/after, and proves the evals themselves by perturbation — the mutation check's sibling: an eval no prompt-degradation can fail is vacuous and gets deleted.
Arguments
[mode] — light (default) | full | plan (preflight, print the plan block, STOP).
[target] — a prompt/instruction file, an agent directory, or changed (prompt-bearing files changed vs HEAD, default).
--ask — pause at the preflight plan for confirmation; default is no gate — the plan prints and the run starts (references/preflight.md).
Modes
Tier vocabulary and platform mechanics: references/model-tiers.md.