composite-scene
Installation
SKILL.md
Composite scene
Combine several real elements (a product, a subject, a backdrop, a style) into one coherent image, with lighting and perspective already reconciled, without cutting out, masking, or compositing by hand. Each element is a separate reference image and the prompt directs how the pieces fit. This is the reverse of character-consistency: instead of holding one subject across many images, it pulls many images into one.
Inputs to collect
- One clean reference per element. The product shot, the subject, the backdrop, and/or the style reference. A subject on a plain background composites more predictably than one already buried in a busy scene. (Ask only if none provided.)
- The relationship between elements: placement, scale, and contact ("resting on", "leaning against", "walking beside"). The references can't convey this, so the prompt must.
- The target scene: setting, target lighting, camera angle/framing.
- How many outputs, and whether the same elements drop into several different scenes.
Models
- Default: Google Nano Banana 2 (
google:4@3) accepts up to 14 reference images in one call and reconciles lighting and perspective across all of them. Best general pick for composition. - Reference-guided alternatives: any image model that accepts
referenceImages(e.g. Nano Banana Pro, IP-Adapter on a FLUX/SDXL base). Smaller reference budgets and weaker cross-element relighting, so confirm support and the field viarunware-modelsplusrunware-runbefore calling. - Confirm the model is
liveand current viarunware-models. Never hardcode a stale choice.