compare-models-blindly
Installation
SKILL.md
Compare Models Blindly
Compare usable outputs without revealing model identity, teacher answers, latency, or prior scores to the judge.
Freeze the comparison set
Select cases by predefined metadata and stable input order before inspecting outcomes. Record the manifest. Reuse saved inference when available; do not rerun a model merely to prepare judging.
Validate candidates independently before strategic or qualitative judging. A candidate that fails identifier, replay, terminal-boundary, or other safety gates cannot win.
Build the blinded packet
Input rows should contain shared case context and two candidate fields. Run: