ollama
Installation
SKILL.md
Ollama — run open-weight LLMs on one box
Ollama serves GGUF models from a local daemon at http://localhost:11434, exposing both a native
HTTP API and an OpenAI-compatible layer. Your job: reach for the right command, the right endpoint,
and the right quant for the hardware in front of you — and recognize when the model does not fit
and the work belongs on a remote GPU instead.
This skill owns: install/serve, pull/tag, the local API (native + OpenAI-compat), Modelfiles, quantization choice, and VRAM/RAM sizing on a single machine.
When to use / when not
Use when the model runs on this machine: pulling/running a model, fixing an OOM, choosing
Q4 vs Q8, authoring a Modelfile, or wiring an app to localhost:11434.
Go elsewhere when: