seedance
Seedance 2.0 — Prompting & Capabilities
Seedance 2.0 (ByteDance Seed team, launched Feb 2026) is a multimodal audio-video model: it takes text, images, audio, and video in a single call and generates a 4–15s clip with native synchronized audio (dialogue + lip-sync, sound effects, ambient, music) in the same pass.
Core mental model — this drives every decision below: Seedance 2.0 is a physics and cinematography simulator you direct like a film crew, not a mood board you brainstorm with adjectives. It renders physical interactions ("tires smoke as the car pivots on wet asphalt") and directorial language ("slow dolly-in", "golden-hour rim light"). It does not render vague adjectives ("cinematic", "epic", "amazing"). Write shot directions, not vibes.
Capabilities at a glance
| Mode | Input | Use it for |
|---|---|---|
| Text-to-video (T2V) | prompt only | Generating a scene from scratch |
| Image-to-video (I2V) | 1 start image + prompt | Animating an existing still |
| First-and-last frame | 2 images (first_frame+last_frame) + prompt |
Controlled transitions / reveals / before-after |
| Reference-to-video (R2V) | 1–9 images + ≤3 videos + ≤3 audio (max 12 files) + prompt | Character/style/motion/object consistency, voice, "omni" composition |
| Multi-shot | prompt with a shot list / timeline | A short narrative (up to ~6 shots) in one 15s generation |
| Video-extend | prior clip + prompt | Continuing past 15s by chaining |
Native specs: 480p/720p native (1080p and native 4K on hosting platforms), 4–15s, 24 fps, aspect ratios auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Audio is on by default (generate_audio: true), at no extra cost. Full specs, access routes, model IDs, pricing, and the raw API request shape are in references/access-and-specs.md — read it when the user asks how to call/access the model or needs exact limits.