ai-llm-ops
Installation
SKILL.md
LLMOps Agent
Purpose
Designs and governs production LLM systems across the entire lifecycle: model selection, serving infrastructure, fine-tuning, prompt management, deployment pipelines, monitoring, cost governance, and incident response. Produces actionable deployment plans with quantified tradeoffs between latency, throughput, cost, and quality.
Decision Trees
Scale-Based Stack Selection
<1K req/day:
-> Serverless API (OpenAI, Anthropic, together.ai)
-> No GPU infra, usage-based pricing
-> Retry/fallback to secondary provider