What a bilingual voice agent actually costs to run
Vendors quote build cost and go quiet on running cost. Here is the full picture — per-minute economics, the hidden line items, and the break-even calculation you should be doing before you sign anything.
Almost every conversation we have about voice agents starts in the wrong place. The prospective client asks what it costs to build. We answer, and then we ask what they currently spend to answer a phone call. Almost nobody knows.
That second number is the one that decides whether any of this is worth doing, and the fact that it is so rarely calculated is why so many voice deployments end up being justified on vibes.
This is our attempt to put the whole cost picture in one place — including the parts that are inconvenient for us to publish.
The build is the smaller number
A scoped voice agent build starts around AED 60,000 with us and rises with integration complexity. That is a real number and it is the one people fixate on.
Over a three-year horizon, though, running cost usually exceeds it — often by a wide margin. A business handling 2,000 calls a month at an average of three minutes is buying 72,000 minutes a year. At any realistic per-minute rate, that dwarfs the build fee by year two.
This matters because it changes what you should be negotiating. Arguing the build fee down by fifteen percent is worth far less than getting the per-minute economics and the containment rate right.
The line items
Here is what actually appears on the bill, including the ones that tend to be left out of proposals.
Speech and model usage. The dominant cost. You are paying for speech-to-text, the language model reasoning over the conversation, and text-to-speech, usually bundled into a per-minute rate. Bilingual handling costs more than monolingual, because switching languages mid-sentence means running detection continuously rather than once.
Telephony. Your carrier charges for the call regardless of who answers it. This is not new spend — you pay it today — but it belongs in the model so the comparison is honest.
Failed and abandoned calls. People hang up. Calls fail to connect. Some callers immediately ask for a human. You are billed for the seconds consumed either way. In our deployments this is typically five to twelve percent of billed minutes and it is almost never in a vendor’s estimate.
Human escalation. Every call the agent hands over still costs you a person, plus the seconds the agent spent before handing over. An agent with a forty percent containment rate is not saving you forty percent of your call-handling cost; it is saving you rather less, because the other sixty percent now costs slightly more than it did before.
Tuning. This is the one that surprises people. In month one you will review transcripts weekly and change things. That is a real cost, whether it is your team’s time or a retainer. It tapers, but it does not go to zero — your prices change, your services change, and the agent has to change with them.
Hosting and orchestration. Modest. This is not compute-heavy infrastructure and it usually rounds to a rounding error next to the model usage.
The calculation you should actually do
Work out your fully-loaded cost per handled call today. Not the salary — the salary plus employer costs, plus the share of supervision, plus the desk, divided by the number of calls that person actually handles in a month. For most Gulf operations this lands somewhere that surprises the person doing the sum for the first time.
Then compare that against the agent’s cost per contained call — that is, cost per minute multiplied by average handle time, divided by your containment rate. The division by containment rate is the step people skip, and it is the step that matters most.
The output of that comparison is one of three answers:
- The agent is clearly cheaper per contained call, and volume is high enough that the saving pays back the build inside a sensible period. Build it.
- The agent is cheaper per call but volume is too low for the saving to repay the build in under two years. Do not build it yet. Revisit when volume grows.
- The agent is not cheaper per call. Either the containment rate assumption is wrong, or the use case is not a good fit for automation. Find out which before spending anything.
Where the threshold usually falls
For a single-channel, single-language agent doing a well-defined job — booking, order status, qualification — the maths tends to start working somewhere around 400 to 500 calls a month. Below that, the build cost has too few calls to amortise across.
Bilingual raises the threshold, because both the build and the per-minute cost go up. Multi-system integration raises it again.
There is an important exception. If the calls you are missing are high-value — a test drive enquiry, a private-hire booking, a new patient — the calculation is not about cost saving at all. It is about recovered revenue, and a handful of recovered deals a month can justify the whole thing at volumes far below the cost-saving threshold. Those two business cases are completely different and should never be blended into one spreadsheet.
The things that make it worse than modelled
Average handle time creeping up. An agent that is too chatty burns minutes. We have seen a thirty-second increase in average handle time wipe out a meaningful share of the projected saving. Concision is a cost control, not a style preference.
Containment measured optimistically. If you count “the caller hung up” as contained, your containment rate is fiction. Measure resolution, not termination.
Scope creep after go-live. The agent works, so someone asks it to also handle complaints, or refunds, or a second brand. Each addition has a build cost and a tuning cost, and they are rarely budgeted because the original project is “done”.
The things that make it better than modelled
Calls you are currently not answering at all. If thirty percent of your after-hours calls go to voicemail and half of those never call back, the agent is not replacing a cost — it is recovering revenue that currently does not exist in any spreadsheet. This is frequently the largest single item and the one most often left out.
Consistency. Every caller gets the same price and the same information. In operations where quoting varies by whoever answered, the margin recovered from that consistency alone can exceed the cost saving.
Peak absorption. You stop staffing for the peak. The value here is not the headcount removed; it is the calls you were losing at peak that you now capture.
What we do about it
We build this model with you during scoping, before quoting, using your actual volumes and your actual cost per call. Sometimes the output is that you should not build anything yet, and we would rather deliver that in week one than in month nine.
If you want to run the numbers on your own operation, bring a month of call logs to a readiness call and we will do it with you. It takes about twenty minutes and it is free, because a client who understands their own economics is a considerably better client than one who does not.