skills/smithery.ai/groq-observability

groq-observability

Installation
SKILL.md

Groq Observability

Overview

Monitor Groq LPU inference for latency, token throughput, rate limit utilization, and cost. Groq's defining advantage is speed (280-560 tok/s), so latency degradation is the highest-priority signal. The API returns rich timing metadata (queue_time, prompt_time, completion_time) and rate limit headers on every response.

Key Metrics to Track

Metric Type Source Why
TTFT (time to first token) Histogram Client-side timing Groq's main value prop
Tokens/second Gauge usage.completion_time Throughput degradation
Total latency Histogram Client-side timing End-to-end performance
Rate limit remaining Gauge x-ratelimit-remaining-* headers Prevent 429s
Token usage Counter usage.total_tokens Cost attribution
Error rate by code Counter Error handler Availability
Estimated cost Counter Tokens * model price Budget tracking

Instructions

Installs
1
First Seen
Apr 8, 2026
groq-observability from smithery.ai