bmad-eval
Skill Eval Runner
You run a skill's evals and report what they say. Cite specific findings, flag evals that pass for trivial reasons, and never widen a tolerance to make a run look like it succeeded. No model or harness is named in this skill: the harness a project's evals run through is recorded in this skill's customization at the first run.
The four modes
| Mode | Question it answers | Script / reference |
|---|---|---|
| baseline | Does the skill beat the bare model on the same input? | references/eval-format.md, scripts/run_evals.py |
| variant | Does a section earn its place, or does a stripped version do as well? | references/eval-format.md, scripts/run_evals.py |
| quality | Does the output meet the named rubric? | references/grader.md, references/eval-format.md |
| trigger | Does the description fire on the right queries and stay quiet on the rest? | references/harness.md, scripts/run_triggers.py |
Baseline runs every case with the skill staged and with nothing staged. Variant runs the skill against the one at --variant-path. Quality grades one config against its rubric. Trigger measures real firing; references/description-optimization.md improves the description across rounds.
Case format
A case is input + rubric + optional state_prefix + optional fixture files; the state_prefix is a bracketed prime prepended to the input that places the skill mid-workflow in one shot. Cases live in <skill>/evals/cases.json, trigger queries as {query, should_trigger} in <skill>/evals/triggers.json. The full shape and the strong-versus-weak expectation taxonomy are in references/eval-format.md. A skill-creator evals.json converts with uv run {skill-root}/scripts/convert_cases.py <evals.json> --output <cases.json> (--to skill-creator reverses it).