jev-eval

Installation
SKILL.md

Evaluating a typed-decision config

The eval set is the product. A System One model's accuracy is dominated by how the question was written, and the failure mode is silent — it returns a confident, type-valid, wrong answer. Without labels you cannot tell a bad question from a bad model.

Measured: rewriting the criteria moved an open model from 4/15 to 14/15 on 15 records. No model change. Then the same comparison at 150 records put that model at 61% overall against Jev's 97% — the 15-record read was an artifact of a small, easy set. Both facts are the point: wording swings results, and small sets lie about which way.

Run it

python ~/.claude/skills/jev-eval/scripts/sweep.py labelled.json configs.json \
    --backend jev|von --question <name>
Installs
7
GitHub Stars
310
First Seen
8 days ago
jev-eval — onewave-ai/claude-skills