higgsfield-audio

Installation
SKILL.md

Higgsfield Audio Prompting Guide

QUICK FACTS

Routing aids — read the linked sections for the full rules.

  • Native-joint audio models: Kling 3.0, Seedance 2.0 / 1.5 Pro, Veo 3/3.1, Grok — all others add audio in post
  • Four layers to consider per prompt: Dialogue / SFX / Ambient / BGM
  • Lip-sync is the most failure-prone feature: 3–8s clips, MCU framing, one speaking face, locked camera, no head-motion tokens; per-language sync-word budgets are FIELD-reported
  • Seedance 2.0 @Audio1 is a conditioning INPUT — beat sync, the [AUDIO: Xs] script block, and the first-15s extraction trap
  • Scope an audio reference like an image one: name the property that rides, the property that must NOT, and where the excluded one comes from instead
  • Multi-clip assembly: one master track · cuts land on musical punctuation, never inside a sung vowel (ECU mouth-match is the one exception) · unified grain + LUT masks batch color drift
  • Cinema Studio 3.0 native joint audio (SCELA): describe audio as a separate section; specific foley beats generic moods
  • Seed Audio 1.0 (seed_audio, standalone) = whole-scene audio in ONE pass — multi-speaker dialogue + music + SFX + ambience mixed
  • Standalone Audio catalog (2026-08-01 snapshot): seed_audio, qwen_audio_tts (NEW — Qwen 3.0 TTS Flash, expressive instructions + cloned voices), text2speech_v2 (5 engines incl. cozy_voice), plus 3 game-pipeline-only tools — distinct from in-video joint audio

Which Models Support Audio?

Installs
12
GitHub Stars
519
First Seen
Apr 14, 2026
higgsfield-audio — osidemedia/higgsfield-ai-prompt-skill