agent-eval
Installation
SKILL.md
Measure agent quality you can defend and gate on
Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.
Do NOT use — route instead
| The ask | Route to | Why it is not this skill |
|---|---|---|
| Build the agent loop, tools, RAG plumbing | building-agents |
It builds the system; you score it. They cross-link. |
| "Make the answers shorter / rewrite the prompt" | prompt-engineering |
Evals say it is worse; that skill changes the words. You never edit the prompt. |
| pytest/jest on deterministic functions | testing-py / testing-web |
Assert-equals on pure code, not stochastic outputs scored by a judge. |
| Dashboards / tracing of live production traffic | observability |
Online monitoring; you are offline + pre-merge. |
| Red-team, jailbreak, prompt injection | agent-safety |
Adversarial coverage, not quality measurement. |
| Per-token cost budgets and accounting | cost-tracking |
You report cost-per-task as one metric; the discipline lives there. |
| A/B stats on product/funnel metrics | ab-testing |
Web experiments, not offline model comparison on a fixed set. |
The eval anatomy
Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.