skills/smithery.ai/groq-cost-tuning

groq-cost-tuning

Installation
SKILL.md

Groq Cost Tuning

Overview

Optimize Groq inference costs through smart model routing, token minimization, and caching. Groq pricing is already extremely competitive, but at high volume the savings from routing classification to 8B vs 70B are 12x per request.

Groq Pricing (per million tokens)

Model Input Output
llama-3.1-8b-instant ~$0.05 ~$0.08
llama-3.3-70b-versatile ~$0.59 ~$0.79
llama-3.3-70b-specdec ~$0.59 ~$0.99
meta-llama/llama-4-scout-17b-16e-instruct ~$0.11 ~$0.34
whisper-large-v3-turbo ~$0.04/hr —

Check current pricing at groq.com/pricing.

Instructions

Installs
1
First Seen
Mar 31, 2026
groq-cost-tuning from smithery.ai