speech-to-text
Installation
SKILL.md
Deepgram Speech-to-Text
Deepgram transcribes audio with two model families on two endpoints. Pick the family first; the endpoint, the parameters, and the message shapes all follow from that choice. This skill gets a first request working and names the skill to open next. It does not repeat the full parameter reference; that lives in the api skill.
Pick the model family first
| Nova | Flux STT | |
|---|---|---|
| Model names | nova-3 (alias of nova-3-general), nova-3-medical, nova-3-pharma |
flux-general-en (English), flux-general-multi (10 languages) |
| Endpoint | /v1/listen, REST and WebSocket |
/v2/listen, WebSocket only |
| Output | A transcript stream | TurnInfo events carrying turn state and a transcript per turn |
| Turn detection | None built in; you use endpointing and your own logic | Built in: StartOfTurn, EagerEndOfTurn, TurnResumed, EndOfTurn |
| Formatting and analysis | smart_format, diarize_model, summarize, sentiment, topics, intents, redaction |
Word timestamps, numerals, redact (numbers or aggressive_numbers), keyterm, profanity_filter, mip_opt_out, tag; no smart formatting, no diarization |
| Language | language=<code>, or language=multi for code-switching |
The model name selects the language; language_hint biases flux-general-multi |
Decision rule: