Why per-minute cost transparency matters in AI calling
A single blended rate hides where your money goes and takes away your ability to control it. Component-level pricing puts you back in charge of cost per call.

Most AI calling platforms sell you a single number: so many rupees per minute. It sounds simple, and simplicity is comforting when you are trying to budget a campaign. But that single number hides a stack of moving parts, and the opacity is not an accident. A blended per-minute rate lets a vendor blend a fat margin into a figure you cannot decompose, and it stops you from making the one optimization that matters most at scale: matching the cost of each component to the value of each call.
Under the hood, every AI voice minute is really three services stitched together plus some add-ons. Speech-to-text converts what the customer says into text. A language model reads that text, along with your instructions and any data it looks up, and decides what to say and do. Text-to-speech turns the reply into spoken audio. On top of that sit telephony, orchestration, and any tool calls the agent makes. Each of these has its own cost curve, and they are wildly unequal.
This piece is a deep look at that cost stack: what each component really costs, why text-to-speech usually dominates, why blended rates quietly hurt you, and how to match a model stack to a use case so you are not overpaying on simple calls or underperforming on hard ones. It ends with worked examples in rupees and dollars, and the reselling math for partners who want to build a margin on top.
The four layers of a call cost
Break any AI voice minute into its parts and you get four layers. Understanding them individually is the whole game, because you cannot optimize a number you cannot see inside.
- Speech-to-text (STT): transcribes the customer's speech in real time. Usually the cheapest of the three AI components, but quality matters enormously for Indian accents and code-switching.
- Language model (LLM): the brain that interprets intent, follows your instructions, and decides the next action. Cost scales with how much context you feed it per turn and which model tier you pick.
- Text-to-speech (TTS): converts the reply into natural audio. Premium neural voices are compute-heavy per second and frequently the single largest line item in the call.
- Add-ons: telephony minutes, orchestration, tool calls to your systems, and any follow-up attempts. Individually small, but they add up across volume.
The reason this matters is that these layers are not close to equal, and the biggest one is not the one people assume. Most buyers assume the LLM is the expensive part because it is the impressive part. In practice, on a typical conversational call, the voice you hear is often the biggest cost driver.
Why text-to-speech usually dominates
Text-to-speech is billed by how much audio it generates, and a voice agent talks a lot. Every greeting, every clarifying question, every readback of a phone number or amount, every reassurance is synthesized speech. Premium neural voices, the ones that sound genuinely human and pronounce Indian names correctly, cost meaningfully more per second than robotic older voices. Multiply that per-second cost by the many seconds the agent speaks across a call, and TTS often ends up as the largest component.
This has a direct, practical implication. If you want to reduce cost on a high-volume campaign, the highest-leverage lever is frequently the voice, not the language model. A reminder bot that reads a mostly fixed script does not need the most expensive premium voice; a good mid-tier voice can sound perfectly professional at a fraction of the per-second cost. Across hundreds of thousands of minutes, that single choice can move your unit economics more than any other.
The corollary is that verbosity is expensive. An agent that over-explains, repeats itself, and delivers long monologues is burning TTS on every extra word. Tightening the script so the agent says what is needed and then listens is both better customer experience and lower cost. When you can see the component breakdown, you can actually notice that your TTS bill is high because your prompt makes the agent talk too much.
Why blended rates hurt you
A blended per-minute rate collapses all four layers plus margin into one figure. That feels convenient, but it causes three specific harms. First, you cannot tell how much of the number is service cost and how much is vendor margin, so you have no basis to negotiate or compare. Second, you cannot optimize, because the lever that would cut your cost, choosing a cheaper voice on simple calls, is invisible and often not even offered. Third, the price becomes unpredictable: when you add a feature or the vendor silently upgrades a model, the number moves and you cannot explain why.
Blended pricing also encourages a one-size-fits-all stack. If the vendor charges everyone the same rate, they will run everyone on the same models, which means either they overprovision the simple calls or underprovision the hard ones. Neither serves you. Transparent, component-level pricing, by contrast, lets you assemble the right stack for each agent and see exactly what it costs, starting from a sensible baseline of around 5.83 rupees per minute and moving up or down as you change components.
There is also a trust dimension. A vendor willing to show you the component costs is telling you they are comfortable being measured on the merits. A vendor who guards a single opaque number is usually protecting a margin they would rather you not see. In a market where you will run this system for years and scale it across campaigns, that transparency compounds into real money and real leverage.
Matching the model stack to the use case
The core insight of transparent pricing is that different calls deserve different stacks. Spending premium-voice, premium-LLM money on a two-line appointment reminder is waste. Cutting corners on a nuanced retention or hardship conversation is false economy. Once you can choose STT, LLM, and TTS per agent, you can right-size each one.
Think of your calls on a spectrum from scripted to reasoning-heavy. Scripted calls, such as reminders and simple confirmations, follow a narrow path with predictable branches. Reasoning-heavy calls, such as objection handling, hardship negotiation, or complex support, require the agent to understand nuance and adapt. The right stack shifts along that spectrum.
- Scripted, high-volume (reminders, COD confirmation, simple surveys): a fast lower-cost LLM, an efficient STT, and a solid mid-tier voice. Optimize hard for cost because volume is huge and the conversation is narrow.
- Semi-structured (lead qualification, appointment booking, first-level support): a mid-tier LLM that can handle branching, a strong STT for accuracy, and a good voice. Balance cost and quality.
- Reasoning-heavy (retention, hardship collections, complex troubleshooting): a stronger LLM justified by the value of the outcome, premium STT for accuracy on emotional or fast speech, and a premium voice because trust matters. Here quality earns its keep.
The discipline is to ask, for each agent, what the call is worth and how hard the conversation is, then provision accordingly. A platform that lets you do this per agent turns cost from a fixed tax into a dial you control.
Budgeting and forecasting
To budget an AI calling program, you need three inputs: the number of contacts you will attempt, the average talk time per connected call, and your connect rate. Attempts multiplied by connect rate gives connected calls; connected calls multiplied by average minutes gives billable minutes; minutes multiplied by your per-minute stack cost gives spend. The trap is forgetting that unconnected attempts still consume small telephony and dialing costs, and that follow-up attempts multiply the contact count.
A robust forecast models the whole funnel, not just the connected minute. If you attempt a list three times to reach the people who did not pick up the first time, your total attempts can be double or triple your contact list. Each of those attempts carries some cost even when nobody answers. Build this into the model so your actual bill does not surprise you.
- Estimate attempts including retries: a 3-attempt cadence on a list of contacts can mean two to three times as many dial attempts as contacts.
- Apply a realistic connect rate: connect rates vary widely by time of day, number reputation, and audience, so use a conservative figure and refine with real data.
- Multiply connected calls by average talk minutes: reminders may average under a minute; support or negotiation calls can run several minutes.
- Add the tail: unanswered-attempt costs, tool calls, and follow-up messages. Small per unit, real in aggregate.
Unit economics: cost per outcome, not cost per minute
Per-minute cost is an input, not the metric that should drive decisions. What matters is cost per outcome: cost per confirmed order, cost per captured promise-to-pay, cost per qualified lead, cost per resolved query. A slightly more expensive stack that resolves more calls can have a lower cost per outcome than a cheap stack that fumbles conversations and forces re-contact.
To compute it, divide total campaign spend by the number of successful outcomes. This reframes the whole cost conversation. Suddenly a premium voice that lifts your promise-to-pay rate is not an expense, it is an investment that lowers cost per successful outcome. And an overly verbose agent that inflates minutes without improving results is exposed as pure waste. Always tie cost back to the business result the calls exist to produce.
This is also how you compare AI calling to the alternative. A human agent in a domestic BPO costs far more per productive minute once you load in salary, attrition, training, supervision, and idle time. When you express AI calling as cost per outcome and compare it to the human cost per outcome, the economics usually become obvious, especially on high-volume, repetitive calls.
Worked cost examples
Here are illustrative examples to make the stack concrete. These are for reasoning, not quotes; your real numbers depend on your chosen models and telephony. Assume a transparent baseline stack around 5.83 rupees per minute, roughly 7 US cents at an indicative exchange rate, for a balanced mid-tier configuration.
- Example A, reminder campaign: 100,000 attempts, 40 percent connect, average 40 seconds of talk time. That is 40,000 connected calls at about 0.67 minutes each, roughly 26,700 billable minutes. On a cost-optimized cheaper-voice stack of around 5 rupees per minute, that is roughly 1.33 lakh rupees, or about 1,600 dollars, plus a small tail for unanswered attempts.
- Example B, lead qualification: 20,000 attempts, 35 percent connect, average 2.5 minutes. That is 7,000 connected calls and about 17,500 billable minutes. On a balanced stack around 5.83 rupees per minute, roughly 1.02 lakh rupees, or about 1,225 dollars. If this yields, say, 1,400 qualified leads, cost per qualified lead is around 73 rupees, which you compare against your cost of a qualified lead through other channels.
- Example C, hardship collections call: 5,000 attempts, 30 percent connect, average 3.5 minutes. That is 1,500 connected calls and about 5,250 minutes. On a premium stack, say 7 rupees per minute, roughly 36,750 rupees, or about 440 dollars. If this drives incremental recovered EMIs worth many multiples of that, cost per rupee collected is what actually matters, and it will look very favorable.
Notice how the right stack differs across the three. The reminder optimizes for the cheapest sensible voice because volume dominates. The collections call justifies a premium stack because the outcome value is high and nuance matters. Blended pricing would have forced the same rate on all three and hidden these choices from you entirely.
Reselling margins for partners
If you are an agency or software vendor reselling voice AI to your own clients, transparent component pricing is not just nice, it is the basis of your business model. When the underlying cost is visible and low, you have room to set a client-facing price that includes a healthy margin while still being fair to your client. When the underlying cost is a single opaque number, you are reselling someone else's mystery and cannot defend your pricing or protect your margin.
The reselling math is straightforward once costs are transparent. Take the true stack cost per minute, add your operating cost to manage the client, and add your target margin. Because the base can start around 5.83 rupees per minute, a partner can price to clients at a rate that is still attractive versus human calling while retaining a substantial per-minute margin. White-label branding lets you present this as your own product, so the client sees your brand and your price, not the underlying platform.
- Know your true cost: the transparent component stack per minute is your floor.
- Add your service layer: onboarding, prompt tuning, QA, and support have real cost; price them in.
- Set margin deliberately: a transparent low base gives room for meaningful per-minute margin while staying below human-calling cost for the client.
- Protect it with white-label: your brand, your pricing, your client relationship, with the underlying costs invisible to the client but fully visible to you.
Takeaways
A single blended per-minute rate is a convenience that costs you money. Every AI voice minute is really speech-to-text, a language model, and text-to-speech plus telephony and orchestration, and those layers are deeply unequal, with premium voice frequently the largest. Seeing the components is what lets you optimize: right-size the stack per use case, tighten verbose prompts, and stop overpaying on simple calls.
Budget the whole funnel including retries, judge the program on cost per outcome rather than cost per minute, and, if you resell, build your margin on a transparent floor with white-label branding. Transparency from around 5.83 rupees per minute is not just a pricing style, it is the mechanism that turns cost from a fixed tax into a dial you actually control.
Written by Callaro Team
The team building Callaro — outcome-driven AI voice agents for teams that live on the phone.
Related articles
How much do AI voice agents cost? A 2026 per-minute breakdown
The three pricing models for AI voice agents, what makes up a per-minute cost, and the real all-in ranges in 2026 — so you can budget without surprises.
Best AI voice agent platforms in 2026: an honest comparison
How the major AI voice agent platforms actually differ in 2026 — by pricing model, model choice, languages and who they're built for.
See the platform run a real call.
Book a demo and we'll show you outcomes executed end-to-end.
Book a demo