self-hosted
Installation
SKILL.md
Deepgram Self-Hosted
Deepgram ships its inference stack as container images you run on your own NVIDIA GPUs. Audio never leaves your network; only small license-verification messages do. This skill decides whether you need that, gets you licensed, and routes to the deployment target.
Self-hosting is gated on an Enterprise agreement. Read Decide before you deploy first — a large share of "we need on-prem" requirements are actually data-residency requirements that a regional endpoint already satisfies, with none of the GPU operations.
Decide before you deploy
Read down the table and take the first row whose requirement you actually have. If none of them apply, the hosted API is the answer.
| Requirement | Answer | What you operate |
|---|---|---|
| Audio may not leave your network, or you need bare metal, an air gap, or FIPS crypto | Full self-hosted containers | GPUs, drivers, containers, models, scaling |
| Must run inside your AWS account, and you want AWS to handle provisioning and scaling | Amazon SageMaker via AWS Marketplace | A SageMaker Endpoint |
| Dedicated capacity, regional control, or compliance isolation — but not your hardware | Deepgram Dedicated — {SHORT_UID}.{REGION}.api.deepgram.com |
Nothing — existing API keys work |
| Data must be processed in the EU, Australia, or India | Regional endpoint: api.eu.deepgram.com, api.au.deepgram.com, api.in.deepgram.com |
Nothing — same API keys and SDKs, change the base URL only |
| None of the above | Hosted API, api.deepgram.com |
Nothing |
People reach for the top row first. Most requirements are satisfied by a row further down.