ai-evaluate
Installation
SKILL.md
Blind pairwise eval — the method
Reproducible procedure for comparing K generator variants — decision hierarchies, workflows, prompt versions, anything that produces an output from an input — on identical frozen inputs. Never self-grade: the rubric author, the variant producers, and the judges are all separate agents, and judges are blind to variant identity. The unit under test can be anything with a name and an output; only the mechanism producing the output is allowed to vary.