Skip to content
CostPerPrompt

Voice AI Cost Calculator

The full three-part bill — speech-to-text + LLM + text-to-speech — priced per minute, per call, and per month. LLM prices updated 2026-08-02.

LLM usage assumes ~1,800 context tokens in and ~220 tokens out per conversation minute — typical for turn-by-turn voice agents.

Cost per minute $—
Per call
$—
Per month
$—
STT / LLM / TTS split
vs human agent ($0.75/min)

Reference prices used (checked 2026-08-02)

LayerServiceApprox. price
Speech-to-textOpenAI Whisper API$0.006/min
Speech-to-textOpenAI GPT-4o Transcribe$0.006/min
Speech-to-textDeepgram Nova$0.0043/min
Speech-to-textAssemblyAI Universal$0.0062/min
Speech-to-textGoogle Cloud STT$0.016/min
Text-to-speechOpenAI TTS — ~$15/1M chars ≈ $0.015/min of speech$0.015/min
Text-to-speechGoogle Cloud TTS (Neural)$0.016/min
Text-to-speechElevenLabs (Creator tier) — premium quality$0.1/min
Text-to-speechCartesia Sonic$0.02/min

LLM cost is computed live from our model pricing data. For the LLM layer alone across every model, use the API cost calculator.

Frequently asked questions

How much does a voice AI agent cost per minute?

A typical stack in 2026 lands between $0.03 and $0.15 per conversation minute: speech-to-text ($0.004–$0.016/min) + LLM reasoning ($0.005–$0.08/min depending on model and verbosity) + text-to-speech ($0.015–$0.10/min depending on voice quality). Premium voices (ElevenLabs-class) can push the total past $0.15/min — at scale, TTS choice is usually the biggest lever.

What dominates the bill: STT, LLM, or TTS?

For flagship-LLM stacks, the LLM dominates. For mini-model stacks (the sensible default for routine calls), premium TTS becomes the largest line item — often 50–70% of the total. Cheap-TTS + mini-LLM stacks get a full voice minute under $0.04.

How do these per-minute costs compare to human agents?

A human call-center minute costs $0.50–$1.50 in most Western markets (fully loaded). Even a premium voice AI stack at $0.15/min is 3–10× cheaper, which is why voice agents are among the fastest-growing AI workloads — and why getting the per-minute math right before scaling matters.

What hidden costs should I budget for?

Telephony (Twilio-class per-minute charges, ~$0.007–$0.02/min), silence and hold time (you often pay STT for it), interruption handling (regenerated TTS), and conversation logging/analytics. Add 20–30% to the raw stack estimate for a realistic production budget.