a-b-testing-agent-output
Installation
SKILL.md
A/B Testing Agent Output
A/B testing agent output compares two or more variants of AI-generated content in production to determine which performs better. The test might compare AI vs human output, Prompt A vs Prompt B, Model A vs Model B, or AI-personalized vs template-only. The goal is to make data-driven decisions about when and how to use AI in your GTM workflow.
The principle: AI output quality is measurable, not assumed. "The AI writes good emails" is an opinion. "AI-personalized emails produced 11.2% reply rate vs 7.8% for templates, with 95% confidence across 400 sends" is a test result. Test before scaling. Measure continuously after scaling.
What to A/B Test
The 5 test types for agent output
| Test type | Variant A | Variant B | What you learn |
|---|---|---|---|
| AI vs human | AI-generated email/content | Human-written email/content | Whether AI matches or exceeds human quality |
| Prompt A vs Prompt B | Output from prompt version 1 | Output from prompt version 2 | Which prompt produces better output |
| Model A vs Model B | Output from Claude Sonnet | Output from Claude Opus (or Haiku) | Whether the more expensive model produces measurably better results |
| AI-personalized vs template | AI-generated first line + template body | Template only (no personalization) | Whether AI personalization lifts reply rates enough to justify the effort |
| Full AI vs AI-assisted | Fully AI-written email | AI-generated first line, human-written body | Whether full AI matches the quality of human-AI hybrid |