agentic-evals
Installation
SKILL.md
agentic-evals
Use this skill when a normal one-shot smoke test is too weak and the task needs repeatable, evidence-backed evaluation of a skill or agent workflow.
Current Scope
This initial bundle provides a deterministic fixture runner for command-based cases. It runs each case multiple times, records stdout/stderr/exit status and duration, checks explicit expectations, and emits a machine-readable readiness summary.
This proves only the declared fixture behavior. It does not prove semantic correctness, real service integration, LLM-judge quality, or release readiness unless the fixture commands themselves exercise those live paths.