agent-eval

Installation
SKILL.md

Measure agent quality you can defend and gate on

Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.

Do NOT use — route instead

The ask Route to Why it is not this skill
Build the agent loop, tools, RAG plumbing building-agents It builds the system; you score it. They cross-link.
"Make the answers shorter / rewrite the prompt" prompt-engineering Evals say it is worse; that skill changes the words. You never edit the prompt.
pytest/jest on deterministic functions testing-py / testing-web Assert-equals on pure code, not stochastic outputs scored by a judge.
Dashboards / tracing of live production traffic observability Online monitoring; you are offline + pre-merge.
Red-team, jailbreak, prompt injection agent-safety Adversarial coverage, not quality measurement.
Per-token cost budgets and accounting cost-tracking You report cost-per-task as one metric; the discipline lives there.
A/B stats on product/funnel metrics ab-testing Web experiments, not offline model comparison on a fixed set.

The eval anatomy

Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.

Installs
2
GitHub Stars
116
First Seen
Aug 6, 2026
agent-eval — ericrisco/rsc-harness