voice-agent

Installation
SKILL.md

Deepgram Voice Agent API

Audio in, audio out, one connection. Deepgram runs the listen, think, and speak stages and the turn-taking between them. Your code streams the user's audio, plays the agent's audio, and answers function calls. [1][2]

Decide first: the Voice Agent API or an orchestrator

Both paths are supported; Deepgram publishes guides for LiveKit Agents and Pipecat. Choose by who should own the pipeline.

Build on the Voice Agent API when Use an orchestrator (LiveKit Agents, Pipecat, Vapi, Retell) with Deepgram STT and TTS underneath when
You want one WebSocket and no pipeline code. End-of-turn detection, barge-in, and the handoffs between stages are handled in-process. [2] You already run that framework, or you need its transport (for example WebRTC rooms) and client libraries.
A Deepgram-managed LLM (OpenAI, Anthropic, Google, NVIDIA) billed through your Deepgram account is fine, or you point think.endpoint at your own OpenAI-compatible endpoint. [6] You need per-stage control the agent does not expose: your own LLM loop, a TTS vendor Deepgram does not proxy, custom voice activity detection, or your own turn logic.
Your tools can run in your client or behind an HTTP endpoint you own (FunctionCallRequest / FunctionCallResponse). [9][10] Your tools live inside the framework's agent runtime.

For the orchestrator path, load the examples skill (LiveKit, Pipecat) and the SDK conversational-stt and text-to-speech skills. Deepgram's Pipecat guide runs Flux STT and Flux TTS (flux-alexis-en) underneath. The LiveKit guide runs three stages: LiveKit Inference (deepgram/aura-2 voice thalia, hosted and billed through LiveKit Cloud with no Deepgram API key), the Deepgram plugin with your own key on nova-3 and aura-2-thalia-en, then STTv2 flux-general-en and TTSv2 flux-alexis-en as the Flux STT and Flux TTS swap. [14] The rest of this skill covers the Voice Agent API path.

First request

The agent host is agent.deepgram.com; api.deepgram.com serves the other APIs. This call lists the LLM models Deepgram can run for you; check a think.provider.model value here before it goes into Settings. The endpoint is public: the Authorization header is optional and the call answers 200 without it, so it also confirms the host is reachable before you spend a key on the socket. [6]

Installs
11
Repository
deepgram/skills
GitHub Stars
23
First Seen
14 days ago
voice-agent — deepgram/skills