llm-serving-capacity-planner
Installation
SKILL.md
LLM Serving Capacity Planner
Overview
Use this when a serving log has enough memory lines to explain where GPU HBM went. The analyzer reads SGLang/vLLM startup logs, extracts weight load, KV pool, CUDA graph, framework overhead, and token-capacity lines, then estimates concurrent requests for common token lengths.
Confirmation Required
Before running analysis, collect or verify these inputs: