benchmark

Installation
SKILL.md

# benchmark

A benchmark is only as fair as its blindest part: identical prompts, isolated runs, a judge who doesn't know whose work it's scoring, and stats you actually measured.

1. Setup — quick by default

Parse the argument for a target (model or skill) and a mode (deep; quick otherwise). Ask only what's missing:

  • No target → one question: benchmark models or skills?
  • Model · quick — contenders are whatever the user named, else the session's model plus one sensible rival; one task. Model · deep — ask how many and which; 2–3 tasks.
  • Contenders can be Claude models at any effort (fable, opus, sonnet, haiku × low/medium/high/xhigh/max) or external CLI agents (codex, gemini, cursor-agent, …) — see §3 for how each runs.
  • Skill · quick — one skill, one task, two conditions: with-skill vs bare on the same model. Skill · deep — several skills, or two skills pitted against each other on shared ground.

In quick mode never ask more than one question total — pick a task yourself, state it, and let the user veto.

2. Pick a calibrated task

The user's own task always wins — if they brought one, use it verbatim. Otherwise pick: right-sized means a sub-agent finishes in roughly 2–5 minutes and a few thousand output tokens. Not fizzbuzz (everyone aces it, nothing separates contenders); not "build an app" (slow, expensive, judging turns mushy). Suggest 2–3 matched to the target:

Installs
6
First Seen
Jun 11, 2026
benchmark — h00mankind/workflow