together-kueue
Installation
SKILL.md
Kueue on Together GPU clusters
Kueue is a Kubernetes-native job queueing controller. It holds jobs in a queue and admits them only when their quota is free, suspending the rest. Use it to share a fixed GPU pool across teams or workloads without overcommitting it. Unlike Volcano, Kueue does not replace the scheduler; it gates when jobs start by toggling their suspend flag.
Public cookbook: https://docs.together.ai/docs/kueue-on-gpu-clusters. Pair this skill with the together-gpu-clusters skill, which covers creating the cluster and configuring kubectl.
Preconditions
- A Together Kubernetes GPU cluster in the
Readystate, withkubectlpointed at it (tg beta clusters get-credentials <cluster_id> --set-default-context). - GPU nodes expose
nvidia.com/gpu(NVIDIA device plugin preinstalled on Together clusters).
Install
Pin the version. The API group is kueue.x-k8s.io/v1beta2 in v0.18.x; older examples using v1beta1 will not apply.
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v0.18.3/manifests.yaml
kubectl -n kueue-system wait --for=condition=Available deployment/kueue-controller-manager --timeout=240s