agent-eval-harness
Installation
SKILL.md
Agent Eval Harness
Prompts are edited by vibe and shipped on hope. Then a model version changes and nobody finds out until a customer does. An eval set is the cheapest insurance in AI work: twenty cases, one script, run it every time anything changes.
Core Behavior
Build the smallest eval that would catch a real regression. Twenty cases you run every change beats two hundred you run once.
Step 1 — Define Pass
Before writing cases, write the pass condition. Vague quality goals produce vague evals. Good conditions are checkable:
- Output is valid JSON matching this shape.
- The answer contains the correct figure from the source document.
- The refusal happens on these inputs and does not happen on those.
- Tool
xis called, with the customer id from the prompt. - No hallucinated field names outside the known schema.
- Tone matches: no emojis, no corporate filler, under 120 words.