speech-to-text

Installation
SKILL.md

Deepgram Speech-to-Text

Deepgram transcribes audio with two model families on two endpoints. Pick the family first; the endpoint, the parameters, and the message shapes all follow from that choice. This skill gets a first request working and names the skill to open next. It does not repeat the full parameter reference; that lives in the api skill.

Pick the model family first

Nova Flux STT
Model names nova-3 (alias of nova-3-general), nova-3-medical, nova-3-pharma flux-general-en (English), flux-general-multi (10 languages)
Endpoint /v1/listen, REST and WebSocket /v2/listen, WebSocket only
Output A transcript stream TurnInfo events carrying turn state and a transcript per turn
Turn detection None built in; you use endpointing and your own logic Built in: StartOfTurn, EagerEndOfTurn, TurnResumed, EndOfTurn
Formatting and analysis smart_format, diarize_model, summarize, sentiment, topics, intents, redaction Word timestamps, numerals, redact (numbers or aggressive_numbers), keyterm, profanity_filter, mip_opt_out, tag; no smart formatting, no diarization
Language language=<code>, or language=multi for code-switching The model name selects the language; language_hint biases flux-general-multi

Decision rule:

Installs
17
Repository
deepgram/skills
GitHub Stars
24
First Seen
Sep 19, 2026
speech-to-text — deepgram/skills