slow-lane
slow-lane
arc-llm-proxy runs on the operator box (LAN 192.168.1.159, local 127.0.0.1) on port 8091, fronting two upstream endpoints: the model box's llama-server (192.168.1.103:1234, Bonsai-27B) and the Veles cloud GPU (Qwen3.8-27B-GGUF via trycloudflare tunnel). One port, six role aliases — /v1/models is the whole model surface:
| alias | lane | use |
|---|---|---|
planning |
fast | interview, planning, design |
hard |
fast | difficult execution |
easy |
fast | cheap local tasks |
bench |
slow | benchmarks, evals |
driver |
slow | plan execution, orchestration |
hygiene |
slow | nightly self-improvement cron (dream/token-waste/adaptation-review/gap-remediate) |
Slow-lane aliases queue and dispatch only into idle slots, always reserving one slot for fast traffic. #slow/#fast on any alias overrides its lane.
Rule
Long-running or non-user-facing LLM calls use a slow-lane alias (bench/driver). If it would wait, let it wait — that is the point. Cron jobs never use claude — they run on these free local models; compensate for model limits with pipeliner designs (small chained steps, explicit retries). Never hit :1234 on the model box directly — you'd bypass the queue and starve interactive requests.