higgsfield-audio
Installation
SKILL.md
Higgsfield Audio Prompting Guide
QUICK FACTS
Routing aids — read the linked sections for the full rules.
- Native-joint audio models: Kling 3.0, Seedance 2.0 / 1.5 Pro, Veo 3/3.1, Grok — all others add audio in post →
- Four layers to consider per prompt: Dialogue / SFX / Ambient / BGM →
- Lip-sync is the most failure-prone feature: 3–8s clips, MCU framing, one speaking face, locked camera, no head-motion tokens; per-language sync-word budgets are FIELD-reported →
- Seedance 2.0
@Audio1is a conditioning INPUT — beat sync, the[AUDIO: Xs]script block, and the first-15s extraction trap → - Scope an audio reference like an image one: name the property that rides, the property that must NOT, and where the excluded one comes from instead →
- Multi-clip assembly: one master track · cuts land on musical punctuation, never inside a sung vowel (ECU mouth-match is the one exception) · unified grain + LUT masks batch color drift →
- Cinema Studio 3.0 native joint audio (SCELA): describe audio as a separate section; specific foley beats generic moods →
- Seed Audio 1.0 (
seed_audio, standalone) = whole-scene audio in ONE pass — multi-speaker dialogue + music + SFX + ambience mixed → - Standalone Audio catalog (2026-08-01 snapshot):
seed_audio,qwen_audio_tts(NEW — Qwen 3.0 TTS Flash, expressive instructions + cloned voices),text2speech_v2(5 engines incl. cozy_voice), plus 3 game-pipeline-only tools — distinct from in-video joint audio →