google-gemini-omni-video
Google Gemini Omni video
Use this skill for the first-party Google model gemini-omni-flash-preview. Treat all availability, prices, limits, and preview behavior below as verified on 2026-07-10 and re-check the cited first-party pages before a production launch.
Establish scope before acting
Ask which backend, data region, source media, edit mode, aspect ratio, desired duration, audio/dialogue, intended audience, rights status, and spend approval apply. Do not silently translate an Omni request into a different product.
| Surface | Contract in scope | Do not conflate it with |
|---|---|---|
| Gemini Developer API | POST https://generativelanguage.googleapis.com/v1beta/interactions; API/auth key; paid tier only for Omni |
Consumer Gemini app, AI Studio UI behavior, or Google Cloud data controls |
| Gemini Enterprise Agent Platform | POST .../v1beta1/projects/{project}/locations/{location}/interactions; OAuth/IAM; Agent Platform API; exact host depends on verified location contract |
A drop-in Developer API endpoint or the legacy Veo operation schema |
| Gemini Omni Flash | Native short video generation, references, audio generation, and conversational editing | Veo, which is a specialized video family with different endpoints and features such as extension/interpolation in supported Veo versions |
| Gemini native image (“Nano Bananaâ€) | Separate image generation/editing models | Omni video output |
| Gemini app, Flow, YouTube | Consumer/creator products with their own access, settings, and terms | A promise about API retention, schema, availability, or provenance metadata |
| Third-party gateway/reseller | Out of scope for this skill | First-party Google billing, privacy, safety, or retry guarantees |
FACT: Gemini Omni Flash is a preview model, although the Interactions API itself is GA. The Developer API model page specifies text/image/video input, video output, a 1,048,576-token context, and 3–10 second 720p/24 FPS output. The Cloud model page publishes a different 131,072-token input limit and 57,920-token output limit. Apply limits only to the backend that documents them.