google-cloud-speech
Installation
SKILL.md
Google Cloud Speech for media production
Use Google Cloud Speech as a provider option when the job is fundamentally about speech audio:
- Transcribe existing audio/video into scripts, searchable text, edits, subtitles, captions, or speaker notes.
- Generate SRT/WebVTT captions from long media.
- Plan live captions or real-time transcript features.
- Create synthetic narration, dialogue prototypes, educational voiceovers, IVR-style prompts, or localized voice tracks.
- Use Google voice families such as Chirp 3 HD, Gemini-TTS, Studio, Neural2, WaveNet, or Standard when the production brief values Google Cloud governance, scale, regional controls, or existing GCP infrastructure.
- Create a consented custom voice only when official access, voice-owner consent, and rights review are present.
Do not use this skill as a generic audio-editing, music, denoise, DAW, mixing, or video-localization skill. Google Cloud Speech can produce transcripts and voice assets, but production finishing still needs editing, loudness normalization, sync, QC, and often translation tools outside the Speech APIs.
Volatile facts below were verified against official Google Cloud documentation on 2026-07-10. Re-check model IDs, launch stages, regional endpoints, quotas, and prices before committing spend or compliance claims.
API boundary
Google exposes related but separate services: