scenario-gemini-image
Scenario Gemini Image
Overview
Gemini, Google's image family on Scenario (the Nano Banana line), generates and edits through one required prompt: edits are instructions against the references, never mask painting. Discover members with search and treat model_schema_get as the contract: the creative fields are shared and nearly everything else is per member.
Connection and the core loop: see the scenario skill in this repo; model-agnostic image work: the scenario-image skill. Gemini video belongs to the scenario-gemini-omni skill; Gemini speech models are the audio domain. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.
Quick reference
Three members at authoring time (per-member facts move, so read the schema):
| Member | Resolution | Sets it apart |
|---|---|---|
| 3.1 Flash | 512 to 4K, default 1K |
video input, four thinkingLevel steps, Search grounding |
| 3.0 Pro | 1K to 4K, default 2K |
built for complex instruction edits and multi-image fusion |
| 3.1 Lite | fixed 1K, no resolution |
fastest and cheapest; thinkingLevel is MINIMAL or HIGH |
Shared fields: referenceImages (up to 14, an array even for one), numOutputs (1 to 4 variations of one prompt), aspectRatio (21:9 through 9:16 plus the default auto; pin the ratio when a placement demands one), and a prompt cap near 250000 characters, so a full brief fits verbatim. Flash alone takes video (one clip, about 15 MB, sampled at videoFps, default 1 fps) to pull stills from footage; video and referenceImages are mutually exclusive. useGoogleSearch (Flash and Pro) grounds the run in live web context and moves the price, like every field marked cost_impact. No mask, seed, or negative-prompt field exists: regional edits are sentences, and reruns give variations, not reproductions.