Developer-first speech AI platform with real-time speech-to-text, Aura text-to-speech and a unified Voice Agent API. $200 non-expiring free credit,

Overview

Deepgram is speech AI infrastructure rather than a consumer transcription app. It sells APIs: Nova-3 for speech-to-text, Aura-2 for text-to-speech, and a Voice Agent API that bundles listening, LLM orchestration and speaking behind one WebSocket connection. That packaging matters because most teams building phone agents currently stitch three vendors together and inherit three sources of latency. Nova-3 covers 45+ languages with speaker diarization, smart formatting, keyterm prompting and automatic language detection, and Deepgram bills per second rather than rounding audio up to the nearest 15 seconds - a detail that quietly saves 10-20% on short clips. In 2026 the company added Flux, a conversational STT model with built-in turn detection and interruption handling aimed specifically at real-time agents. Deployment options include cloud, VPC and fully self-hosted, which is why Deepgram turns up in healthcare and call-center stacks where audio cannot leave the perimeter. The trade-off is that it gives non-developers nothing out of the box: no meeting-notes UI, no collaborative editor, no shared workspace. You get endpoints, SDKs, a console and generous free credit. If you want a finished product, buy a transcription app; if you are building one, this is the layer underneath it.

Key Features

  • Nova-3 speech-to-text across 45+ languages with diarization and smart formatting
  • Flux conversational STT with native turn detection for real-time voice agents
  • Aura-2 text-to-speech billed per 1,000 characters
  • Unified Voice Agent API combining STT, LLM orchestration and TTS over one socket
  • True per-second billing with no rounding up of short audio
  • Self-hosted and VPC deployment for regulated audio workloads
  • Audio Intelligence add-ons for summarization, topic detection, sentiment and intent

Pricing

PlanPriceFor
Pay As You Go$200 free credit, then from $0.0077/min (Nova-3 mono)Developers and prototypes
Growth$4,000+/year prepaid, up to 20% offProduction apps at scale
Voice Agent API$0.050-$0.163/min depending on BYO componentsReal-time phone and web agents
EnterpriseCustomHigh volume, self-hosted, SLAs

Comparison

Compared to Whisper, Deepgram trades open weights for managed real-time streaming, diarization and an SLA - Whisper is free to self-host but you own the GPU bill and all the latency engineering. Compared to ElevenLabs, the two only overlap on text-to-speech: ElevenLabs wins on voice cloning and expressive delivery, Deepgram wins on cost per character and on keeping transcription in the same account. Vapi sits one layer higher and can call Deepgram as its transcription provider, so the real decision is buy-the-orchestrator versus build-on-the-API.

Compare alternatives

Side-by-side with the 3 closest alternatives.

ToolCategoryPricingVisit
Deepgram (this) voice, audio, codeFrom $0.05/mo Site ↗
Whisperaudio, productivityFrom $0/mo Site ↗
ElevenLabsvoiceFree $0 · From $5/mo Site ↗
Vapivoice, agents, codeFrom $0.05/mo Site ↗
Deepgram Current

Developer-first speech AI platform with real-time speech-to-text, Aura text-to-speech and a unified Voice Agent API. $200 non-expiring free credit,

voiceaudiocode
From $0.05/mo

OpenAI's open-source speech recognition model for accurate multilingual transcription. Free and open-source Read our hands-on review and compare the top

audioproductivity
From $0/mo

Text-to-speech and voice cloning platform with some of the most natural AI voices available. Most natural TTS voices on the market Read our hands-on

voice
Free $0 · From $5/mo

Developer-first voice AI platform for building phone and web voice agents with your choice of STT, LLM and TTS providers.

voiceagentscode
From $0.05/mo
Editor’s Review
4.5/5
Pros
  • +$200 non-expiring free credit, roughly 400+ hours of Nova-3 transcription, no card required
  • +True per-second billing instead of rounding audio up to 15-second or one-minute blocks
  • +A single Voice Agent API removes the latency of stitching separate STT, LLM and TTS vendors together
Cons
  • Developer-only: there is no end-user app, transcript editor or shared workspace
  • Add-ons stack up fast - diarization plus PII redaction can add roughly 50% to the base per-minute rate
  • Growth-tier discounts require a $4,000 annual prepayment

Deepgram is the default choice when transcription is a build dependency rather than a finished product. Accuracy is competitive with Whisper at meaningfully lower operating cost, and per-second billing is an honest touch in a category full of rounding tricks. It is the wrong tool for anyone who wants an app rather than an endpoint.

See all reviews →

Last updated: 2026-08-04

When to use it

  • Use it when you need $200 non-expiring free credit, roughly 400+ hours of Nova-3 transcription, no card required
  • Use it when you need true per-second billing instead of rounding audio up to 15-second or one-minute blocks
  • Use it when you need a single Voice Agent API removes the latency of stitching separate STT, LLM and TTS vendors together

When to skip it

  • Avoid it if developer-only: there is no end-user app, transcript editor or shared workspace
  • Avoid it if add-ons stack up fast - diarization plus PII redaction can add roughly 50% to the base per-minute rate
  • Avoid it if growth-tier discounts require a $4,000 annual prepayment

Alternatives to consider