skill-eval

Installation
SKILL.md

/skill-eval

Author one behavioral probe for one skill, at the cheapest tier that can still separate the arms, and report the verdict honestly. A probe measures behavior-change — did loading the skill change what the agent did — never quality-uplift. This skill authors and tiers probes. scripts/probe-skill.sh runs them.

Insight: when a probe returns INERT because the control arm already aces the scenario, the measurement failed, not the skill. Weakening the producer is one escape and it costs realism. The cheaper escape is to plant the defect: build a scenario containing exactly one flaw the discipline catches and a skim does not, then grade whether the agent acted on it. Signal you manufacture is signal you can reproduce.

The failure mode this exists to prevent: a skill catalog whose tier badges are editorial. A skill nobody measured is a skill nobody can defend, and re-running a saturated scenario at a lower effort level produces more rows in the ledger without producing more knowledge.

Installs
4
Repository
boshu2/agentops
GitHub Stars
431
First Seen
5 days ago
skill-eval — boshu2/agentops