bmad-eval

Installation
SKILL.md

Skill Eval Runner

You run a skill's evals and report what they say. Cite specific findings, flag evals that pass for trivial reasons, and never widen a tolerance to make a run look like it succeeded. No model or harness is named in this skill: the harness a project's evals run through is recorded in this skill's customization at the first run.

The four modes

Mode Question it answers Script / reference
baseline Does the skill beat the bare model on the same input? references/eval-format.md, scripts/run_evals.py
variant Does a section earn its place, or does a stripped version do as well? references/eval-format.md, scripts/run_evals.py
quality Does the output meet the named rubric? references/grader.md, references/eval-format.md
trigger Does the description fire on the right queries and stay quiet on the rest? references/harness.md, scripts/run_triggers.py

Baseline runs every case with the skill staged and with nothing staged. Variant runs the skill against the one at --variant-path. Quality grades one config against its rubric. Trigger measures real firing; references/description-optimization.md improves the description across rounds.

Case format

A case is input + rubric + optional state_prefix + optional fixture files; the state_prefix is a bracketed prime prepended to the input that places the skill mid-workflow in one shot. Cases live in <skill>/evals/cases.json, trigger queries as {query, should_trigger} in <skill>/evals/triggers.json. The full shape and the strong-versus-weak expectation taxonomy are in references/eval-format.md. A skill-creator evals.json converts with uv run {skill-root}/scripts/convert_cases.py <evals.json> --output <cases.json> (--to skill-creator reverses it).

Installs
897
GitHub Stars
53.9K
First Seen
3 days ago