We build voice agents for a living, so you would expect us to say the ROI is spectacular. It is usually good. It is almost never what the vendor decks claim. The pitch most operators hear promises 95% automation and payback in weeks, and that pitch quietly assumes numbers that production call traffic does not deliver. We covered the engineering side in our post on sub-700ms voice agents; this one is the other half: the arithmetic a clinic, a logistics operator, or a services business should run before signing anything.
The short version: the economics work, but they work for different reasons than most buyers expect, and only for well-scoped call types. Here is what is true.
The short answer: an AI-handled phone call costs $0.30–$0.80 all-in versus $2–$4 for a human-handled one. A well-scoped voice agent contains 40–70% of calls in production — not the 95% in vendor decks — and a scoped, integrated build from AGI Software Solutions typically runs $15,000–$35,000 with payback in 4–8 months, with recovered after-hours revenue often exceeding the labor saving in year one.
What a human-handled call actually costs
Start with an honest denominator. A front-desk person or phone CSR does not cost their salary; they cost salary plus benefits, payroll taxes, a seat, software licenses, a slice of a supervisor, and the hiring-and-training cost of the 30–45% annual turnover typical for phone-heavy roles. Fully loaded, the phone-answering hour in most Western markets lands between $20 and $30. In India and similar markets the absolute number is lower, but so is the revenue per call, so the ratios come out surprisingly similar.
Now divide by realistic throughput. Nobody handles calls for 60 minutes of every clocked hour. Between wrap-up notes, callbacks, and the front-desk reality of doing three jobs at once, 8–12 handled calls per hour is a good day. That puts the fully-loaded cost of a human-handled call at roughly $2 to $4. And the part that hurts is not the cost per call; it is that the cost stays fixed while the volume does not. Humans cannot burst. Every spike becomes hold time, and hold time becomes abandonment.
What an AI-handled call actually costs
An AI phone call is a metered pipeline: telephony in, speech-to-text, a language model, text-to-speech out. Realistic per-minute costs as we scope them today:
- Telephony: $0.005–0.02 per minute depending on carrier and country.
- Streaming STT: $0.004–0.01 per minute.
- LLM: $0.01–0.06 per minute, driven mostly by prompt size and how much conversation history you carry.
- TTS: $0.02–0.08 per minute for voices you would actually put in front of customers.
Raw stack: roughly $0.05–0.15 per minute, so a typical three-minute call costs $0.15–0.50 in usage. Add orchestration, monitoring, and a margin for retries and you should model $0.30–0.80 per call all-in. Managed platforms flatten this to $0.10–0.25 per minute and are the right answer at low volume.
So the gross saving per automated call sits somewhere around $2–3.50. The crossover is therefore a volume question. Below roughly 1,000 in-scope calls a month, the saving will not carry a custom build; use an off-the-shelf platform or leave it alone. Past 2,000–3,000 calls a month, especially once you need real integration with your scheduling or order system, custom starts winning on both per-call cost and containment, because it can complete transactions instead of reading FAQs aloud.
Containment rate is the whole game
Containment is the share of calls the agent resolves end-to-end with no human touch. It is the single number that decides whether the project pays back, and it is the number vendors inflate the most. A demo that answers questions is not containment. A call where the bot talked for four minutes and then transferred anyway is negative containment: you paid for the AI minutes, the human minutes, and the caller’s patience.
In production, a well-scoped voice agent contains 40–70% of calls. The 95% figure in the sales deck counts every call the bot spoke on, not the calls it actually finished.
The bands we see in practice: appointment booking and rescheduling 50–70%, order status and delivery windows 60–70%, FAQ and hours-and-directions traffic above 70%, anything open-ended 25–40% and usually not worth automating yet. Two design decisions move these numbers more than the model choice does: how narrowly you scope the call type, and how early the agent detects it is out of its depth.
That second point matters because the non-contained 30–60% are not free. Every failed call must hand off warm, with a context summary, so the caller never repeats themselves. If your handoff is a cold transfer into a generic queue, callers learn to smash zero on the first ring, and containment collapses along with the ROI.
An AI voice agent for customer service: what changes
Everything above applies to a single-site front desk. Running an AI voice agent for customer service at contact-centre scale changes three things, and each one moves the ROI model.
- Containment is measured per intent, not per queue. A blended containment number across a contact centre is close to meaningless, because it averages 70% FAQ traffic against 25% open-ended traffic. Report it by intent or you cannot tell what to fix.
- CSAT usually moves up before it moves down. The common finding is that satisfaction improves on the contained calls, because answering on the first ring with no hold time beats a human after eight minutes of queue music, and falls on badly handed-off calls. The net depends almost entirely on handoff quality, which is a design decision rather than a model capability.
- Average handle time drops on the human calls too. When the agent has already captured identity, intent and account context, the human conversation starts further along. This second-order saving is real and is almost never in the business case.
Set targets accordingly: 40–70% containment on the intents you scoped, CSAT on contained calls at or above your current baseline, and a measurable reduction in handle time on transferred calls. If a vendor proposes a single blended containment target and no CSAT floor, they are proposing a number you cannot manage.
AI in the call centre: the market claim and the counterpoint
The figure doing the rounds is roughly $80B in contact-centre labour costs addressable by AI in 2026, which is the number underneath most AI call center pitches. It is directionally plausible and close to useless for planning, because it describes what is theoretically automatable rather than what is deployed.
The counterpoint is the number worth holding onto: only about a quarter of contact centres have voice AI genuinely integrated with their systems of record. The rest are running something that can talk but cannot look up an order, verify an identity or complete a booking. That gap is the entire difference between a deflection tool and an automation programme, and it is why containment figures vary so wildly between deployments that describe themselves identically.
The practical read: budget for integration, not for the voice layer. The speech pipeline is close to a commodity in 2026. Knowing your customer, your entitlements and your inventory is not, and that is where the cost and the containment both come from.
The revenue side almost everyone leaves out
Most ROI models only count labor savings. In the projects we ship, the first win is usually revenue, and it shows up before the savings do.
Count your missed calls: after hours, weekends, lunch rushes, and everything that rings out while both staff are mid-conversation. For appointment-driven businesses we typically find 15–30% of inbound calls go unanswered or hit voicemail, and a large share of those callers never leave a message; they call the next provider on the list. A voice agent answers all of them, at 2 a.m. on a Sunday, on the first ring. If even a modest fraction of recovered calls become bookings, that line frequently exceeds the labor line in year one. Missed calls are not a service problem. They are a sales problem wearing a service costume.
What voice AI is still bad at
We turn down automation on entire call categories, and you should too:
- Emotionally loaded calls. Complaints, billing disputes, a patient scared about a result. The technology can sound sympathetic; it cannot be accountable, and callers can tell.
- Multi-issue conversations. A caller who wants to reschedule, ask about insurance, and dispute a charge in one call will defeat most agents. Humans juggle threads; current agents drop them.
- Heavy accents and bad audio beyond what your STT tolerates: warehouse noise, speakerphone in a truck cab, a crowded lobby. Word error rate degrades quietly and containment falls off a cliff.
- Calls where empathy is the product. Bereavement, bad-news delivery, retention saves. If the caller needs to feel heard by a person, an agent that merely simulates it damages the brand.
The design rule: automate around these categories, not through them. Detect them in the first exchange and route to a human fast. A voice agent that knows what it should not handle outperforms one that tries to handle everything.
A worked example: a mid-size clinic front desk
The numbers below are deliberately conservative and match the shape of clinic projects we scope. Swap in your own volumes.
- Three front-desk staff, about 2,400 inbound calls a month, averaging 3.5 minutes.
- About 300 calls a month land after hours or ring out at peak. Historically half of those callers try again; the other ~150 are simply gone.
- Roughly 60% of volume is bookings, reschedules, confirmations, and routine questions. The agent takes this slice and contains 55% of it.
Cost side. 1,440 in-scope calls, roughly 790 contained. At 5–8 minutes of staff time saved per contained call including wrap-up, that is 65–105 hours a month, about half a full-time seat. You will not cut headcount, and you should not plan to; that capacity absorbs growth, kills hold time, and finally gets the no-show reminder calls made, which quietly recovers more revenue than anything else on this page.
Revenue side. Of the ~150 truly lost calls a month, assume 40% carried booking intent and the agent converts half of what it answers. That is about 30 incremental bookings. At $80–150 of revenue per visit, call it $2,500–4,000 a month that currently goes to whichever competitor picks up the phone.
Run cost and payback. 1,440 agent calls at three minutes and ~$0.10 per minute is under $500 a month; budget $600–900 with platform and monitoring. A scoped, integrated build of this kind usually lands in the $15,000–35,000 range in our experience. Net monthly benefit of $3,500–5,500 against that build puts payback at 4–8 months, with numbers we would defend in front of a CFO. The vendor deck says weeks. The truth is months. Months is still an excellent answer.
What to measure from day one
ROI you cannot measure is ROI you cannot defend at renewal time. Instrument these before launch, not after:
- True containment: resolved end-to-end and verified against the system of record (the booking actually exists), not the bot’s own claim of success.
- Transfer quality: on handoffs, did the human get context, and did the caller repeat themselves? Sample and score weekly.
- Booking completion rate: bookings finished versus bookings started. A wide gap means the flow is broken, not the model.
- Caller sentiment: scored per call from the transcript. Watch the trend, not individual calls.
- Latency: response gaps beyond about a second read as incompetence and train callers to demand a human.
We wrote up the full harness in our post on voice agent evals; the short version is that transcript accuracy is the least interesting metric on the board.
Where to start
- Pick exactly one call type. The highest-volume, lowest-emotion one. Appointment booking and order status are the usual winners.
- Measure four weeks of baseline first. Call volume by hour, missed-call rate, minutes per call, booking conversion. Without a baseline, nobody can prove the agent did anything.
- Pilot with a graceful exit. Route a fraction of traffic, keep the handoff warm, and set a containment floor below which you pause and fix rather than push on.
- Decide with a rule, not a vibe. Under ~1,000 in-scope calls a month: use a platform or wait. Above that, with 40%+ projected containment and a real missed-call problem, the arithmetic in this post starts working for you.
If you want a second pair of eyes on your numbers, this is the scoping exercise we run at the start of every voice AI development engagement: your call data, the arithmetic above, and an honest answer about whether the project pays for itself, including the times the answer is not yet.
Frequently asked questions
How much does a voice AI agent cost to build?
A scoped, integrated build — one call type, wired into your scheduling or order system, with warm handoff — usually lands in the $15,000–$35,000 range. Run cost is roughly $0.30–$0.80 per call for a custom stack, or $0.10–$0.25 per minute on managed platforms, which are the right answer below ~1,000 in-scope calls a month.
How much does an AI phone call cost versus a human?
Human-handled: roughly $2–$4 fully loaded, at a realistic 8–12 handled calls per hour. AI-handled: $0.30–$0.80 all-in for a typical three-minute call. And the AI absorbs volume spikes without hold time, which humans cannot.
What containment rate is realistic?
40–70% for a well-scoped agent: bookings 50–70%, order status 60–70%, FAQ traffic above 70%, open-ended calls 25–40% and usually not worth automating yet. The 95% in the sales deck counts every call the bot spoke on, not the calls it finished.
How long until a voice agent pays back?
For an appointment-driven business with 2,000+ monthly calls: 4–8 months, from roughly half a seat of recovered staff time plus recovered missed-call revenue. Weeks-long payback claims assume containment production traffic does not deliver.
Which calls should not be automated?
Emotionally loaded calls, multi-issue conversations, audio beyond what your STT tolerates, and calls where empathy is the product. Detect them in the first exchange and route to a human fast.