ollama

Installation
SKILL.md

Ollama — run open-weight LLMs on one box

Ollama serves GGUF models from a local daemon at http://localhost:11434, exposing both a native HTTP API and an OpenAI-compatible layer. Your job: reach for the right command, the right endpoint, and the right quant for the hardware in front of you — and recognize when the model does not fit and the work belongs on a remote GPU instead.

This skill owns: install/serve, pull/tag, the local API (native + OpenAI-compat), Modelfiles, quantization choice, and VRAM/RAM sizing on a single machine.

When to use / when not

Use when the model runs on this machine: pulling/running a model, fixing an OOM, choosing Q4 vs Q8, authoring a Modelfile, or wiring an app to localhost:11434.

Go elsewhere when:

Installs
3
GitHub Stars
116
First Seen
Aug 6, 2026
ollama — ericrisco/rsc-harness