cost-aware-llm-pipeline

Installation
SKILL.md

Cost-Aware LLM Pipeline

Patterns for controlling LLM API costs while maintaining quality. Combines model routing, budget tracking, retry logic, and prompt caching into a composable pipeline.

When to Activate

  • Building applications that call LLM APIs (Claude, GPT, etc.)
  • Processing batches of items with varying complexity
  • Need to stay within a budget for API spend
  • Optimizing cost without sacrificing quality on complex tasks
  • Designing a multi-model pipeline where simple classification tasks should use Haiku and complex reasoning tasks should escalate to Sonnet or Opus automatically
  • Adding a hard budget cap to a batch processing job so it fails fast rather than silently overspending when processing hundreds or thousands of files
  • Implementing prompt caching for a system prompt that is longer than 1024 tokens and is repeated on every API call in a high-volume pipeline
  • Auditing an existing LLM integration that currently uses the most expensive model for all requests regardless of task complexity

Core Concepts

1. Model Routing by Task Complexity

Installs
2
GitHub Stars
14
First Seen
Apr 7, 2026
cost-aware-llm-pipeline — marvinrichter/clarc