a-b-testing-agent-output

Installation
SKILL.md

A/B Testing Agent Output

A/B testing agent output compares two or more variants of AI-generated content in production to determine which performs better. The test might compare AI vs human output, Prompt A vs Prompt B, Model A vs Model B, or AI-personalized vs template-only. The goal is to make data-driven decisions about when and how to use AI in your GTM workflow.

The principle: AI output quality is measurable, not assumed. "The AI writes good emails" is an opinion. "AI-personalized emails produced 11.2% reply rate vs 7.8% for templates, with 95% confidence across 400 sends" is a test result. Test before scaling. Measure continuously after scaling.

What to A/B Test

The 5 test types for agent output

Test type Variant A Variant B What you learn
AI vs human AI-generated email/content Human-written email/content Whether AI matches or exceeds human quality
Prompt A vs Prompt B Output from prompt version 1 Output from prompt version 2 Which prompt produces better output
Model A vs Model B Output from Claude Sonnet Output from Claude Opus (or Haiku) Whether the more expensive model produces measurably better results
AI-personalized vs template AI-generated first line + template body Template only (no personalization) Whether AI personalization lifts reply rates enough to justify the effort
Full AI vs AI-assisted Fully AI-written email AI-generated first line, human-written body Whether full AI matches the quality of human-AI hybrid

Prioritizing tests

Installs
1
GitHub Stars
1
First Seen
Jun 6, 2026
a-b-testing-agent-output — pfoy/growth-skills