llm-cost-latency-budget

Installation
SKILL.md

LLM Cost & Latency Budget Skill

LLM features have a unit cost and a tail latency that demos hide and production exposes. This skill does the token math up front — what one request costs, what a million cost, where the p95 latency comes from — and lays out the levers (model tiering, caching, prompt trimming) so cost and speed are designed, not discovered.

Required Inputs

Ask for these only if they aren't already provided:

  • The request shape — typical system prompt, user input, retrieved context, and output sizes (in rough tokens).
  • Volume — requests/day now and at target scale; peak concurrency.
  • Models in play — candidate model(s) and their per-token input/output prices.
  • Targets — acceptable cost per request (or per user/month) and the latency users will tolerate (p50 / p95).

Output Format

Cost & Latency Budget: [feature]

Installs
6
GitHub Stars
1.2K
First Seen
Jun 25, 2026
llm-cost-latency-budget — mohitagw15856/pm-claude-skills