Deepgram Voice Agent API pricing looks simple at first because the main unit is a per-minute WebSocket session. For a real support workflow, though, the bill usually depends on where the bundled Voice Agent API stops and where separate TTS, LLM, telephony, browser transport, storage, monitoring, and review costs begin. This cost matrix separates those layers so teams can compare Deepgram, AssemblyAI, OpenAI Realtime, and Makefun-adjacent avatar workflows with the same scenario assumptions.
Quick pricing snapshot
Checked on May 30, 2026, the official Deepgram pricing page lists Voice Agent API pay-as-you-go tiers calculated from WebSocket connection time: Standard at $0.075/min, Standard BYO TTS at $0.065/min, Custom BYO LLM at $0.056/min, Custom BYO LLM + TTS at $0.050/min, Advanced at $0.163/min, and Advanced BYO TTS at $0.122/min. Deepgram also lists Aura-2 text to speech at $0.030 per 1,000 characters and Aura-1 at $0.0150 per 1,000 characters.
Those numbers are the starting point, not the full operating cost. The Deepgram Voice Agent docs describe a single WebSocket flow for realtime conversations, while the LLM model docs separate standard, advanced, and bring-your-own model choices. The rate-limit docs also matter because concurrency, regional endpoint, and Aura-2 capacity shape how far a pilot can scale before commercial limits or architecture changes appear.
Cost matrix by workflow
| Workflow | Deepgram line items | External line items to budget | Best fit check |
|---|---|---|---|
| Browser support voice chat | Voice Agent API session minutes; standard or advanced tier; optional BYO LLM/TTS mode. | Frontend voice UI, browser transport, analytics, transcript storage, QA review, and escalation handling. | Use when fast realtime turn-taking matters and the browser experience is already controlled. |
| PSTN appointment calls | Voice Agent API minutes plus selected STT/TTS/LLM mode. | Telephony provider minutes, SIP/PSTN routing, call recording, consent, failed-call retries, and human handoff. | Use the same call-duration sample for Deepgram, AssemblyAI, and OpenAI comparisons. |
| Multilingual IVR | Flux or Nova STT choices, Aura-2 output, advanced LLM tier if needed, and WebSocket duration. | Locale QA, pronunciation tuning, fallback prompts, regional endpoints, compliance review, and agent transfer rules. | Check language mix before assuming one per-minute rate covers every market. |
| Avatar support handoff | Deepgram conversation layer, Aura-2 or BYO TTS, and tool/function-call flow. | Avatar rendering, video scene generation, lip-sync, asset hosting, moderation, and review. Makefun workflows can handle the visual layer when the support agent needs a talking-avatar output. | Use when voice support should become reusable video or avatar content. |
| Regulated or self-hosted support | Deepgram managed or custom deployment terms, selected model tier, and project limits. | Retention policy, private cloud or on-prem review, audit logging, BAA/security review, and internal support operations. | Do not treat self-hosting, compliance, or enterprise capacity as self-serve pricing. |
Competitor comparison boundaries
AssemblyAI pricing lists its Voice Agent API at $4.50/hr, or $0.075/min, which makes it a useful flat-rate same-use comparator. The comparison is not just price matching: rate limits, included stack behavior, LLM routing, resumption, concurrency, and enterprise controls can change the operating result.
OpenAI Realtime should be compared as a token and audio-token workflow, not as a simple per-minute bundle. The OpenAI Realtime cost guide is useful for explaining how audio input/output, context growth, truncation, cache behavior, and conversation design affect spend. That makes OpenAI stronger as a flexible realtime model comparison, while Deepgram and AssemblyAI are easier to model as voice-agent minute bundles.
Hidden cost checklist
- WebSocket connection duration, including silence, waiting time, and interrupted turns.
- Standard versus advanced LLM tier, or BYO LLM cost when the model is billed elsewhere.
- BYO TTS, Aura-2 character volume, voice selection, pronunciation tuning, and multilingual QA.
- Telephony, SIP, browser transport, recording, storage, transcript retention, and observability.
- Tool calls, retries, transfers, fallback prompts, failed sessions, and post-call review.
- Concurrency limits, regional endpoint choices, enterprise capacity, privacy review, and compliance approvals.
Pilot template
Start with one measured call sample instead of a vendor-wide assumption. For each scenario, record total connected minutes, speaking minutes, silence, interruptions, tool calls, transfers, recording size, transcript retention, and human QA time. Then calculate four versions: Deepgram bundled Voice Agent API, Deepgram component mode with BYO LLM or TTS, AssemblyAI Voice Agent API, and an OpenAI Realtime stack.
For Makefun-adjacent teams, add a fifth line when the voice agent needs to become a visual support asset. A realtime call may later need an avatar explainer, localized talking video, or creator-facing walkthrough. In that case, compare the voice-agent layer with existing Makefun resources such as the GPT-Realtime voice-agent guide, Gemini Live session calculator, Cartesia voice-agent cost matrix, LiveKit Agents calculator, Pipecat Cloud hosting calculator, and Makefun AI Avatar API.
Decision notes
Choose Deepgram Voice Agent API when the team wants a bundled realtime voice-agent path and can model cost from connection minutes. Consider BYO TTS or BYO LLM when the voice, model, compliance, or routing layer already exists elsewhere. Compare AssemblyAI when a flat hourly voice-agent quote is easier for finance teams. Compare OpenAI Realtime when the application needs flexible multimodal model behavior and the team can manage token, audio-token, and context costs carefully.
The practical rule is simple: do not compare only the headline minute rate. Compare the full workflow boundary, including the parts that happen before the call starts and after the transcript is saved.
FAQ
Is Deepgram Voice Agent API priced by spoken audio or connected time?
Deepgram’s pricing page describes Voice Agent API tiers as calculated from WebSocket connection time. That means silence, waiting, interruptions, and long handoff paths should be included in pilot math.
Does Aura-2 replace the full Voice Agent API cost?
No. Aura-2 is the text-to-speech layer priced per 1,000 characters. It may be part of a Deepgram voice-agent stack or a BYO component plan, but it is not the same thing as the complete Voice Agent API session price.
Can this be compared directly with AssemblyAI?
Yes, but only with the same call scenario. AssemblyAI’s Voice Agent API pricing gives a useful flat-rate comparator, while Deepgram’s tier choices and BYO modes require a billing-boundary check.
When should Makefun enter the workflow?
Use Makefun when the voice-agent output needs a reusable avatar, talking-video, creator-facing demo, or visual support workflow rather than only a live call transcript.



