groq-cost-tuning
Installation
SKILL.md
Groq Cost Tuning
Overview
Optimize Groq inference costs through smart model routing, token minimization, and caching. Groq pricing is already extremely competitive, but at high volume the savings from routing classification to 8B vs 70B are 12x per request.
Groq Pricing (per million tokens)
| Model | Input | Output |
|---|---|---|
llama-3.1-8b-instant |
~$0.05 | ~$0.08 |
llama-3.3-70b-versatile |
~$0.59 | ~$0.79 |
llama-3.3-70b-specdec |
~$0.59 | ~$0.99 |
meta-llama/llama-4-scout-17b-16e-instruct |
~$0.11 | ~$0.34 |
whisper-large-v3-turbo |
~$0.04/hr | — |
Check current pricing at groq.com/pricing.