agent-comparison

Installation
SKILL.md

Agent Comparison Skill

Compare agent variants through controlled A/B benchmarks. Runs identical tasks on both agents, grades output quality with domain-specific checklists, and reports total session token cost to a working solution. This skill is exclusively for agent variant comparison — use agent-evaluation for single-agent assessment, and skill-eval for skill testing.

Reference Loading Table

Signal Load These Files Why
selecting benchmark tasks and directory layout (Phase 1) benchmark-tasks.md Loads detailed guidance from benchmark-tasks.md.
example-driven tasks, errors examples-and-errors.md Loads detailed guidance from examples-and-errors.md.
scoring solutions: 5-criteria rubric and effective cost calculation grading-rubric.md Loads detailed guidance from grading-rubric.md.
deciding when to run comparisons; December 2024 baseline data methodology.md Loads detailed guidance from methodology.md.
configuring autoresearch: targets, task formats, eval isolation modes optimization-guide.md Loads detailed guidance from optimization-guide.md.
executing Phase 5 OPTIMIZE step by step optimize-phase.md Loads detailed guidance from optimize-phase.md.
writing the Phase 4 comparison report report-template.md Loads detailed guidance from report-template.md.

Instructions

See references/examples-and-errors.md for error handling. See references/optimize-phase.md for Phase 5 OPTIMIZE full procedure. See references/methodology.md for December 2024 benchmark data.

Installs
13
GitHub Stars
416
First Seen
Mar 23, 2026
agent-comparison — notque/vexjoy-agent