omcustom:agent-eval-framework
Installation
SKILL.md
Agent Eval Framework
Purpose
Provides quantitative, trajectory-based evaluation for oh-my-customcode agents. Fills the measurement gap not covered by existing skills:
| Skill | Coverage | Gap |
|---|---|---|
| harness-eval | SE benchmark task quality (15 tasks) | No efficiency metrics |
| evaluator-optimizer | Qualitative rubric loop | No quantitative gate |
| deep-verify | Release quality (structural/correctness) | No step/latency ratios |
| multi-model-verification | Code correctness across models | No trajectory comparison |
This skill adds efficiency measurement — not just "did the agent succeed?" but "how efficiently did it succeed relative to an ideal trajectory?"
The 4-Metric Framework
Derived from LangChain's deep agent evaluation methodology.