benchmark
Installation
SKILL.md
# benchmark
A benchmark is only as fair as its blindest part: identical prompts, isolated runs, a judge who doesn't know whose work it's scoring, and stats you actually measured.
1. Setup — quick by default
Parse the argument for a target (model or skill) and a mode (deep; quick otherwise). Ask only what's missing:
- No target → one question: benchmark models or skills?
- Model · quick — contenders are whatever the user named, else the session's model plus one sensible rival; one task. Model · deep — ask how many and which; 2–3 tasks.
- Contenders can be Claude models at any effort (fable, opus, sonnet, haiku × low/medium/high/xhigh/max) or external CLI agents (codex, gemini, cursor-agent, …) — see §3 for how each runs.
- Skill · quick — one skill, one task, two conditions: with-skill vs bare on the same model. Skill · deep — several skills, or two skills pitted against each other on shared ground.
In quick mode never ask more than one question total — pick a task yourself, state it, and let the user veto.
2. Pick a calibrated task
The user's own task always wins — if they brought one, use it verbatim. Otherwise pick: right-sized means a sub-agent finishes in roughly 2–5 minutes and a few thousand output tokens. Not fizzbuzz (everyone aces it, nothing separates contenders); not "build an app" (slow, expensive, judging turns mushy). Suggest 2–3 matched to the target: