instruction-eval
Installation
SKILL.md
Condition A/B
Change one condition an agent runs under, run the same prompts before and after, and show the difference.
The conditions surrounding an agent have no verification. You can read code and tests will catch a regression, but a few lines added to instructions or a reference doc dropped in a directory only ever get judged on whether they sound reasonable. Even the person who put them there has no idea whether they change behavior. This skill replaces that guess with an observation.
This runs on Claude Code. Both arms execute as claude -p subprocesses, so the
CLI has to be available.
Who writes what
The report holds content from two sources, visually separated in the HTML. Never hand-write what the script produces, since transcribing only introduces errors.