nvidia-cosmos-video
Installation
SKILL.md
NVIDIA Cosmos video
Use Cosmos as a world-generation and world-simulation system, not as proof of physics. Choose the task, artifact, and runtime together. Names, schemas, licenses, parameter counts, and hardware claims are surface-specific.
This guide was verified against first-party NVIDIA repositories, model cards, NIM documentation, licenses, and technical reports on 2026-07-10. Recheck every linked model card, release note, support matrix, and embedded license immediately before downloading, deploying, post-training, or serving.
Route by family and task
| Family/artifact | What it does | Inputs and outputs | Status/boundary |
|---|---|---|---|
Cosmos 3 Generator (Cosmos3-Nano, Cosmos3-Super) |
General omnimodal world generation; text/image to video, and on supported open runtimes video, sound, and action modes | Full open checkpoint cards describe text/image/video/audio/action input and image/video/audio/action output | Newest open family, released 2026-05-31. Use Generator for media; the separate Reasoner produces text reasoning, not video |
| Cosmos3-Super-Image2Video | Specialized image-to-video checkpoint | One image + text → video | Do not assume Nano/Super and this specialization are interchangeable |
| Cosmos-Predict2.5 2B/14B | Future-world prediction, unified Text2World/Image2World/Video2World; specialized AV/robot variants | Text plus optional first image/video; some variants consume actions or multiview inputs | Current Predict2.5 repository; 2B and 14B pre-trained/post-trained bases, distilled and domain checkpoints have different capabilities |
| Cosmos-Transfer2.5-2B | World-to-world transformation with structural control | Source RGB video + prompt + edge/depth/segmentation/visual-blur controls → transformed video | Current control model; use for Sim2Real/Real2Real structure preservation, not unconstrained text-to-video |
| Predict2 / Predict1 / Transfer1 | Earlier families | Varying text/image/video/control contracts | Predict2 repository is archived and explicitly recommends migration to Predict2.5. Predict1 remains available in NIM. No public universal EOL date was found; call these earlier/legacy, not unsupported unless the exact artifact says so |
| Cosmos Reason, Embed, Tokenizer, Curator, Evaluator | Reasoning, embeddings, compression/data/evaluation | Not a video-generation endpoint | Do not route a generation request to these merely because they carry the Cosmos name |