claude-api-cost-optimizer
Installation
SKILL.md
Claude API Cost Optimizer
Cut Claude API costs by 70–90% using intelligent model selection, caching, and batching.
Quick Start
- Audit your current API calls — identify which tasks use Opus or Sonnet that could use Haiku. Model selection alone saves 10–18x on simple tasks.
- Pick the cheapest model tier for each task: Haiku (cheapest) → Sonnet (mid) → Opus (most expensive, use sparingly). See
references/pricing.mdfor current rates. - Enable prompt caching for repeated context (system prompts, codebases) by adding
"cache_control": {"type": "ephemeral"}to message blocks - Implement cost reporting — track
input_tokens,output_tokens, and cache metrics from API responses
Key Concepts
- Model selection — Haiku for simple tasks (formatting, comments) — cheapest tier. Sonnet for medium (refactoring, debugging) — mid tier. Opus for complex only (architecture, security) — most expensive, use sparingly. See
references/pricing.mdfor current rates. - Prompt caching — Cache large static content (system prompts, codebase context). Cache reads cost 90% less; writes pay off after 1–2 reuses.
- Batching — Combine multiple requests into one API call to eliminate per-request overhead. 80% fewer calls ≈ 80% lower cost.
- Local caching — Cache identical responses locally to skip redundant API calls entirely.
- Context extraction — Send only relevant snippets, not whole files. Smaller inputs = lower costs.
- max_tokens discipline — Set realistic limits; unused token budget is wasted money.