voice-agent
Deepgram Voice Agent API
Audio in, audio out, one connection. Deepgram runs the listen, think, and speak stages and the turn-taking between them. Your code streams the user's audio, plays the agent's audio, and answers function calls. [1][2]
Decide first: the Voice Agent API or an orchestrator
Both paths are supported; Deepgram publishes guides for LiveKit Agents and Pipecat. Choose by who should own the pipeline.
| Build on the Voice Agent API when | Use an orchestrator (LiveKit Agents, Pipecat, Vapi, Retell) with Deepgram STT and TTS underneath when |
|---|---|
| You want one WebSocket and no pipeline code. End-of-turn detection, barge-in, and the handoffs between stages are handled in-process. [2] | You already run that framework, or you need its transport (for example WebRTC rooms) and client libraries. |
A Deepgram-managed LLM (OpenAI, Anthropic, Google, NVIDIA) billed through your Deepgram account is fine, or you point think.endpoint at your own OpenAI-compatible endpoint. [6] |
You need per-stage control the agent does not expose: your own LLM loop, a TTS vendor Deepgram does not proxy, custom voice activity detection, or your own turn logic. |
Your tools can run in your client or behind an HTTP endpoint you own (FunctionCallRequest / FunctionCallResponse). [9][10] |
Your tools live inside the framework's agent runtime. |
For the orchestrator path, load the examples skill (LiveKit, Pipecat) and the SDK conversational-stt and text-to-speech skills. Deepgram's Pipecat guide runs Flux STT and Flux TTS (flux-alexis-en) underneath. The LiveKit guide runs three stages: LiveKit Inference (deepgram/aura-2 voice thalia, hosted and billed through LiveKit Cloud with no Deepgram API key), the Deepgram plugin with your own key on nova-3 and aura-2-thalia-en, then STTv2 flux-general-en and TTSv2 flux-alexis-en as the Flux STT and Flux TTS swap. [14] The rest of this skill covers the Voice Agent API path.
First request
The agent host is agent.deepgram.com; api.deepgram.com serves the other APIs. This call lists the LLM models Deepgram can run for you; check a think.provider.model value here before it goes into Settings. The endpoint is public: the Authorization header is optional and the call answers 200 without it, so it also confirms the host is reachable before you spend a key on the socket. [6]