h3-long-video

Installation
SKILL.md

MiniMax H3 — long videos: a chain of clips with consistency

H3 generates 4–15 seconds per pass. A long video = a chain of clips, where each next clip is generated with references that carry continuity. Your job in this skill is to help the user plan the chain and write the prompt for each link (use the rules from h3-t2v-prompt and h3-ref-prompt).

What holds consistency (strongest first)

  1. Frame-to-frame seam (the main technique): the last frame of clip N becomes the first-frame anchor of clip N+1. The picture matches pixel-perfect at the seam, and the model develops the scene forward from that frame (I2VA: first-frame anchor → action onset → development → result).
  2. Character references: up to 9 images per generation. Build a "character sheet" once (front, 3/4, full body) and attach it to EVERY clip as <Subject N> sources.
  3. Voice reference: up to 3 audio clips. Cut a clean line of the character's speech from the first clip and attach it as an <Audio N> voice-timbre reference to every later clip with dialogue.
  4. Scene/style reference: a location or style frame as a <Subject N> (environment/style).
  5. Video continuation: the previous clip can be attached as <Video N> with the video continuation task type — H3 natively continues from the end of a source video (note: reference video seconds are billed in the API, and you get at most 3 video references).

Per-generation limits: 9 images + 3 videos + 3 audio, 12 files total; audio only together with an image or video.

The "project bible" — a stable text block

The model does not remember past generations. Everything that must stay identical gets repeated verbatim in every prompt:

Installs
1
GitHub Stars
429
First Seen
5 days ago
h3-long-video — alesha-pro/tools