replicate
Replicate platform operations
This skill is about how a model runs in production on Replicate — clients, async, deployments,
Cog packaging, webhooks, scaling, and spend. It is the platform-engineering counterpart to image
prompt craft. If the question is what prompt, aspect ratio, or model family produces a good image,
that is replicate-images, not this skill. Here the mental model is: a prediction is a job. You
either wait for it, poll it, or get pinged about it — and where it runs (shared cold pool vs a private
warm deployment) is a cost-and-latency dial you set deliberately.
Pinned facts (verified 2026-06-02): Python client replicate 1.x (latest 1.0.7), Python 3.8+;
JS client replicate on npm; auth via REPLICATE_API_TOKEN. A 2.0.0aN alpha exists on PyPI but
is NOT the default — pin replicate>=1,<2 so a fresh install never silently pulls it.
Decision: how should this model run?
Pick the row by latency tolerance and whether your process can block. Do not default to run() for
everything — a 10-minute job inside a web request will time out and burn a worker.