vastai-core-workflow-b
Installation
SKILL.md
Vast.ai Serverless Endpoint Rollout
Overview
Replace ad hoc multi-instance orchestration with the provider Serverless control plane. Prove a template independently, establish worker and queue bounds from load evidence, then let the workergroup perform a graceful rolling update.
Prerequisites
- Latency, error-rate, queue-time, concurrency, and cost objectives
- Immutable model/template candidate and a separate canary endpoint
- Initial, minimum, maximum, cold-worker, and inactivity policy
Instructions
Step 1: Prove the candidate template
Launch the new model or environment on a non-production endpoint and verify load, readiness, response schema, and representative outputs.