elevenlabs-media
Installation
SKILL.md
ElevenLabs media
Turn a media brief into usable, inspected files or an integration built against the actual ElevenLabs contract. Assume ELEVENLABS_API_KEY is already supplied; check presence without printing it. Never ask the user to paste it into chat.
Choose the media operation
Read the branch needed for the task, not this entire package:
| User goal | Read | Important distinction |
|---|---|---|
| Narration, expressive dialogue, voice conversion, sound effects, speech cleanup | Speech and audio | TTS, dialogue, STS and isolation use different models and wire formats |
| Transcript, speakers, captions, exact word timing, live captions | Transcription and alignment | Transcribe unknown speech; align an existing transcript; realtime is a separate protocol |
| Original music, lyrics, composition plans, editing, stems, video soundtrack | Music | Prompt, v1 sections and v2/v2.5 chunks are different request contracts |
| Images, image editing, video, talking portraits, reusable references | Images, video and Flows | Dashboard models are not all API models; jobs are asynchronous |
| Translate/dub audio or video, revise a dub, subtitle a translation | Dubbing | Project readiness is not target-language completion; legacy IDs are different |
| Find/design/remix/clone a voice, pronunciation, audiobooks, podcasts, embedded players | Voices and long-form production | Previewing, saving, training and publishing are separate effects |
| WebSocket speech/dialogue, live STT, speech-engine integration | Realtime | Each transport has its own framing, end signals and timing units |
| Authentication, scopes, cost, safe requests, failure recovery | Access and helper | Permission, plan, quota and model failures are not interchangeable |
| What was actually tested, known conflicting docs | Evidence | Authored evals, local tests and live results are separate evidence |