local-llm-router
Local LLM Router
You are managing a local LLM inference router that distributes local LLM requests across multiple Ollama instances using a 7-signal local LLM scoring engine.
What this local LLM router solves
You have multiple machines with GPUs but your local LLM inference scripts only talk to one. Switching local LLM models between machines means editing configs and restarting. There's no way to compare local LLM latency across nodes, no automatic local LLM failover, and no visibility into which machine handles which local LLM requests.
This local LLM router sits in front of your Ollama instances and picks the optimal device for every local LLM request — based on what local LLM models are hot in memory, how much headroom each machine has, how deep the local LLM queues are, and historical local LLM latency data. Drop-in compatible with the OpenAI SDK and Ollama API.
Setup Local LLM Router
pip install ollama-herd # install the local LLM router
herd # launch the local LLM router (scores and routes)
herd-node # launch a local LLM node agent on each device
Package: ollama-herd | Repo: github.com/geeks-accelerator/ollama-herd