vllm
Installation
SKILL.md
vLLM — high-throughput serving of open-weight models
vLLM is the inference engine you put in front of an open-weight model when many requests hit it at
once. Its job — and this skill's — is throughput under concurrency: keep the GPU busy across dozens of
simultaneous requests, not squeeze one prompt out fast. You own the vllm serve flags; the box those
flags run on is a runpod/modal concern.
Why not just loop a transformers generate()? Naive serving runs one request at a time and pads
every batch to the longest sequence, so the GPU idles. vLLM fixes both:
- PagedAttention stores the KV cache in non-contiguous fixed-size blocks (like OS virtual-memory paging), so there is almost no padding/reservation waste and long contexts pack tightly.
- Continuous batching admits and retires requests token-by-token instead of per-batch, so a new request joins the running batch immediately rather than waiting for the slowest one to finish.
Net effect: an order-of-magnitude more concurrent throughput than single-request serving. If you only
ever have one user on a laptop, that machinery is wasted — that is ollama, not this.