Transparent pricing

Voice calls from $0.06/min — you choose the models, you control the cost.

No seat licenses. No bundled black box. Pick the STT, TTS and LLM for each agent and see the exact per-minute cost before a single call goes out.

The entry-level stack

Pick the cheapest core models and voice calls start at about $0.06/min. This is the live per-minute stack you configure in the agent builder, shown component by component before you launch.

  • STT — Soniox Realtime$0.005/min
  • LLM — Gemini 2.5 Flash$0.020/min
  • TTS — Smallest Lightning v3.1$0.035/min

Models & add-ons

Voice pipeline · per agent

Editable

STT

SonioxRealtime

$0.005

LLM

Google Gemini2.5 Flash

$0.020

TTS

Smallest AILightning v3.1

$0.035
Estimated cost / min$0.06

Optional add-ons

Only pay for these when you switch them on. Nothing here is bundled into the base rate.

Advanced noise cancellation

$0.004/min

Cleaner audio on noisy lines and outdoor calls.

Post-call transcription

$0.007/min

Full searchable transcript generated after each call.

LLM prompt overage

$0.004/min

Only applies when your prompt runs over 5,000 tokens. The first 5,000 prompt tokens are included in every call.

Estimated cost per call minute~$0.14/min
  • STT — Soniox Realtime$0.005/min
  • LLM — Gemini 2.5 Flash$0.020/min
  • TTS — ElevenLabs Flash v2.5$0.10/min
Total (live stack)$0.14/min

Premium voice example. Swaps TTS to ElevenLabs Flash v2.5 ($0.10/min) for a more expressive voice. Total lands around $0.14–0.15/min depending on models.

Premium voice example

When the voice is the brand, upgrade the TTS. Swapping in ElevenLabs Flash v2.5 ($0.10/min) — while keeping the same STT and LLM — brings a typical call to around $0.14–0.15/min.

You decide per agent: cost-efficient models for high-volume operational calls, premium voices for customer-facing brand moments.

You choose the models

Three layers, each swappable per agent. Mix and match to hit the accuracy, latency, language coverage and cost that fits the use case.

Speech-to-text (STT)

How the agent hears and transcribes the caller.

SonioxDeepgramElevenLabsAzureAssemblyAI

Text-to-speech (TTS)

The voice your customers hear — from cost-efficient to studio-grade.

Smallest AICartesiaDeepgramElevenLabsGoogle GeminiAzure

Language model (LLM)

The reasoning that runs the conversation and decides the next action.

Google Gemini 2.5 Flashand more

Estimate your savings

Drag the sliders to see monthly Callaro cost at $0.06/min against what the same calls cost with people.

200
4 min
$25/hr

Assumes a fully-loaded US agent cost at 65% productive occupancy (the rest is after-call work, breaks and idle time), over 21 working days/month at $0.06/min. Premium voices and add-ons cost more.

Human cost / month

$10,769

Callaro / month

$1,008

Estimated monthly savings

$9,761

About 91% lower than doing these calls with people.

4,200 calls/month · 16,800 minutes

What's included

Everything below is part of the platform — no per-seat fees, no add-on charges for the dashboard.

No per-seat or per-agent fees — you pay for minutes, not licenses
First 5,000 prompt tokens included in every call
Full dashboards, transcripts, recordings and summaries
Sentiment, outcome tags, QA flags and action logs
Automated follow-ups, retries and multi-touch cadences
Tool calls, CRM writebacks, bookings and payment links
Bring your own telephony — Twilio, Plivo, Telnyx, Vonage, SignalWire, Amazon Connect
REST, Webhooks and WebSockets integrations
Guardrails, warm human handoff and per-agent configuration

Pricing questions

How is the per-minute price calculated?

Every call is billed by the models you choose: one STT rate, one LLM rate and one TTS rate, added together per minute. The entry-level stack (Soniox + Gemini 2.5 Flash + Smallest Lightning v3.1) works out to about $0.06/min. Swap in premium models and the per-minute rate goes up accordingly — you always see the total before you launch.

What counts as an add-on?

Advanced noise cancellation ($0.004/min) and post-call transcription ($0.007/min) are optional and only cost when you enable them. LLM prompt overage ($0.004/min) applies only when a call's prompt exceeds 5,000 tokens — the first 5,000 are included.

Why would I pick a premium voice?

For customer-facing brand calls where the voice matters, you can upgrade TTS — for example to ElevenLabs Flash v2.5 at $0.10/min — bringing a typical call to around $0.14–0.15/min. For high-volume operational calls, the entry-level stack keeps costs low.

Are there seat or user fees?

No. There are no per-seat or per-agent licenses. You pay for the minutes you use, and dashboards, transcripts, follow-ups and integrations are part of the platform.

Do you offer volume or enterprise pricing?

Yes. As your volume grows we work with you on tailored pricing and dedicated support. Talk to sales for a quote built around your projected minutes.

Running high volume or need enterprise controls?

We'll build a quote around your projected minutes.

Talk to sales

See your cost per outcome, live.

Book a demo and we'll configure a stack for your use case and show the exact per-minute price.

Book a demo