llm-serving-capacity-planner

Installation
SKILL.md

LLM Serving Capacity Planner

Overview

Use this when a serving log has enough memory lines to explain where GPU HBM went. The analyzer reads SGLang/vLLM startup logs, extracts weight load, KV pool, CUDA graph, framework overhead, and token-capacity lines, then estimates concurrent requests for common token lengths.

Confirmation Required

Before running analysis, collect or verify these inputs:

Installs
53
GitHub Stars
809
First Seen
May 21, 2026
llm-serving-capacity-planner — bbuf/ai-infra-auto-driven-skills