vllm

Installation
SKILL.md

vLLM — high-throughput serving of open-weight models

vLLM is the inference engine you put in front of an open-weight model when many requests hit it at once. Its job — and this skill's — is throughput under concurrency: keep the GPU busy across dozens of simultaneous requests, not squeeze one prompt out fast. You own the vllm serve flags; the box those flags run on is a runpod/modal concern.

Why not just loop a transformers generate()? Naive serving runs one request at a time and pads every batch to the longest sequence, so the GPU idles. vLLM fixes both:

  • PagedAttention stores the KV cache in non-contiguous fixed-size blocks (like OS virtual-memory paging), so there is almost no padding/reservation waste and long contexts pack tightly.
  • Continuous batching admits and retires requests token-by-token instead of per-batch, so a new request joins the running batch immediately rather than waiting for the slowest one to finish.

Net effect: an order-of-magnitude more concurrent throughput than single-request serving. If you only ever have one user on a laptop, that machinery is wasted — that is ollama, not this.

Version & setup reality (verify at author time — vLLM ships ~weekly)

Installs
2
GitHub Stars
116
First Seen
Aug 6, 2026
vllm — ericrisco/rsc-harness