agent-comparison
Installation
SKILL.md
Agent Comparison Skill
Compare agent variants through controlled A/B benchmarks. Runs identical tasks on both agents, grades output quality with domain-specific checklists, and reports total session token cost to a working solution. This skill is exclusively for agent variant comparison — use agent-evaluation for single-agent assessment, and skill-eval for skill testing.
Reference Loading Table
| Signal | Load These Files | Why |
|---|---|---|
| selecting benchmark tasks and directory layout (Phase 1) | benchmark-tasks.md |
Loads detailed guidance from benchmark-tasks.md. |
| example-driven tasks, errors | examples-and-errors.md |
Loads detailed guidance from examples-and-errors.md. |
| scoring solutions: 5-criteria rubric and effective cost calculation | grading-rubric.md |
Loads detailed guidance from grading-rubric.md. |
| deciding when to run comparisons; December 2024 baseline data | methodology.md |
Loads detailed guidance from methodology.md. |
| configuring autoresearch: targets, task formats, eval isolation modes | optimization-guide.md |
Loads detailed guidance from optimization-guide.md. |
| executing Phase 5 OPTIMIZE step by step | optimize-phase.md |
Loads detailed guidance from optimize-phase.md. |
| writing the Phase 4 comparison report | report-template.md |
Loads detailed guidance from report-template.md. |
Instructions
See
references/examples-and-errors.mdfor error handling. Seereferences/optimize-phase.mdfor Phase 5 OPTIMIZE full procedure. Seereferences/methodology.mdfor December 2024 benchmark data.