agent-eval-harness

Installation
SKILL.md

Agent Eval Harness

Prompts are edited by vibe and shipped on hope. Then a model version changes and nobody finds out until a customer does. An eval set is the cheapest insurance in AI work: twenty cases, one script, run it every time anything changes.

Core Behavior

Build the smallest eval that would catch a real regression. Twenty cases you run every change beats two hundred you run once.

Step 1 — Define Pass

Before writing cases, write the pass condition. Vague quality goals produce vague evals. Good conditions are checkable:

  • Output is valid JSON matching this shape.
  • The answer contains the correct figure from the source document.
  • The refusal happens on these inputs and does not happen on those.
  • Tool x is called, with the customer id from the prompt.
  • No hallucinated field names outside the known schema.
  • Tone matches: no emojis, no corporate filler, under 120 words.
Installs
7
GitHub Stars
310
First Seen
9 days ago
agent-eval-harness — onewave-ai/claude-skills