composite-scene

Installation
SKILL.md

Composite scene

Combine several real elements (a product, a subject, a backdrop, a style) into one coherent image, with lighting and perspective already reconciled, without cutting out, masking, or compositing by hand. Each element is a separate reference image and the prompt directs how the pieces fit. This is the reverse of character-consistency: instead of holding one subject across many images, it pulls many images into one.

Inputs to collect

  • One clean reference per element. The product shot, the subject, the backdrop, and/or the style reference. A subject on a plain background composites more predictably than one already buried in a busy scene. (Ask only if none provided.)
  • The relationship between elements: placement, scale, and contact ("resting on", "leaning against", "walking beside"). The references can't convey this, so the prompt must.
  • The target scene: setting, target lighting, camera angle/framing.
  • How many outputs, and whether the same elements drop into several different scenes.

Models

  • Default: Google Nano Banana 2 (google:4@3) accepts up to 14 reference images in one call and reconciles lighting and perspective across all of them. Best general pick for composition.
  • Reference-guided alternatives: any image model that accepts referenceImages (e.g. Nano Banana Pro, IP-Adapter on a FLUX/SDXL base). Smaller reference budgets and weaker cross-element relighting, so confirm support and the field via runware-models plus runware-run before calling.
  • Confirm the model is live and current via runware-models. Never hardcode a stale choice.

Workflow

Installs
3
GitHub Stars
2
First Seen
Jul 2, 2026
composite-scene — runware/runware-skills