azure-speech
Installation
SKILL.md
Azure Speech for media production agents
Use this skill when Azure Speech in Foundry Tools is a candidate provider for transcription, captions, subtitles, narration, synthetic voices, avatar speech, speech translation, video localization, or production speech QA.
Treat Azure Speech as a family of related services, not one model. Select the narrowest official capability that matches the job, then verify region, quota, language, pricing, and access status before committing.
Facts below were verified against Microsoft/Azure documentation on 2026-07-10. Volatile items such as supported locales, regions, pricing, quota values, API versions, and preview/limited-access status must be rechecked before paid or production use.
First decision: what job is the user actually asking for?
Documented facts:
- Azure Speech exposes speech to text, text to speech, speech translation, custom speech, custom voice, personal voice, text-to-speech avatar, video translation, LLM Speech, Speech SDK, Speech CLI, REST APIs, and containers in different combinations by region and access tier. Sources: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/overview and https://learn.microsoft.com/en-us/azure/ai-services/speech-service/regions
- Speech-to-text modes include real-time transcription, fast transcription, batch transcription, and custom speech. Source: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-to-text
- Text-to-speech modes include real-time synthesis, asynchronous batch synthesis for long-form audio, standard/prebuilt neural voices, HD voices, custom voice, personal voice, and avatar outputs. Sources: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech and https://learn.microsoft.com/en-us/azure/ai-services/speech-service/batch-synthesis
Production routing: