skills/skills.volces.com/claude-api-cost-optimizer

claude-api-cost-optimizer

Installation
SKILL.md

Claude API Cost Optimizer

Cut Claude API costs by 70–90% using intelligent model selection, caching, and batching.

Quick Start

  1. Audit your current API calls — identify which tasks use Opus or Sonnet that could use Haiku. Model selection alone saves 10–18x on simple tasks.
  2. Pick the cheapest model tier for each task: Haiku (cheapest) → Sonnet (mid) → Opus (most expensive, use sparingly). See references/pricing.md for current rates.
  3. Enable prompt caching for repeated context (system prompts, codebases) by adding "cache_control": {"type": "ephemeral"} to message blocks
  4. Implement cost reporting — track input_tokens, output_tokens, and cache metrics from API responses

Key Concepts

  • Model selection — Haiku for simple tasks (formatting, comments) — cheapest tier. Sonnet for medium (refactoring, debugging) — mid tier. Opus for complex only (architecture, security) — most expensive, use sparingly. See references/pricing.md for current rates.
  • Prompt caching — Cache large static content (system prompts, codebase context). Cache reads cost 90% less; writes pay off after 1–2 reuses.
  • Batching — Combine multiple requests into one API call to eliminate per-request overhead. 80% fewer calls ≈ 80% lower cost.
  • Local caching — Cache identical responses locally to skip redundant API calls entirely.
  • Context extraction — Send only relevant snippets, not whole files. Smaller inputs = lower costs.
  • max_tokens discipline — Set realistic limits; unused token budget is wasted money.
Installs
2
First Seen
Apr 23, 2026
claude-api-cost-optimizer from skills.volces.com