jev-eval
Installation
SKILL.md
Evaluating a typed-decision config
The eval set is the product. A System One model's accuracy is dominated by how the question was written, and the failure mode is silent — it returns a confident, type-valid, wrong answer. Without labels you cannot tell a bad question from a bad model.
Measured: rewriting the criteria moved an open model from 4/15 to 14/15 on 15 records. No model change. Then the same comparison at 150 records put that model at 61% overall against Jev's 97% — the 15-record read was an artifact of a small, easy set. Both facts are the point: wording swings results, and small sets lie about which way.
Run it
python ~/.claude/skills/jev-eval/scripts/sweep.py labelled.json configs.json \
--backend jev|von --question <name>