What a voice agent costs to run
Voice is priced by the minute, not by the token, and that single difference changes how you plan it. A well-built agent on Indian telephony costs roughly 2 to 4 rupees a minute to run, all-in. Here is where that goes.
| Stage | Illustrative cost per minute |
|---|---|
| Speech to text | about $0.006 |
| Language model | about $0.007, varies with how much context you send |
| Text to speech | about $0.02 while speaking, so roughly half a call |
| Telephony | 30 to 70 paise on Indian numbers |
| All in | roughly 2.5 rupees a minute |
At 500 three-minute calls a month that is about 4,000 rupees of usage. At 2,000 calls, about 15,000. Prices move, so treat the method as the point rather than the figures, and note what the shape tells you: text to speech is usually the largest line, which is why choosing a cheaper voice for routine confirmations and a better one for the opening is a real cost lever rather than a detail.
Platform products commonly charge 8 to 30 rupees a minute. That is the same pipeline with a margin on it, and at low volume it is good value. The arithmetic reverses somewhere in the low thousands of calls a month.
How a voice pipeline works
Three models in a row, each adding delay. If you run them one after another and wait for each to finish, the caller hears a pause long enough to say "hello?" into. Every stage has to stream: transcription emits words as they are spoken, the model starts generating before the caller finishes, and speech synthesis begins before the sentence is complete.
Latency is the whole product
Callers forgive a wrong answer more readily than a long silence. A wrong answer sounds like a person who is mistaken; a two-second gap sounds like a machine, and once someone decides they are talking to a machine, they stop cooperating with it.
So we budget latency by stage rather than measuring it at the end, and we build interruption handling in from the start. A caller who talks over the agent must be heard immediately, which means the agent has to stop speaking, discard what it was about to say, and listen. Retrofitting that is far harder than designing for it, because it touches every stage of the pipeline at once.
Indian languages, and the code-switching problem
Provider marketing lists dozens of supported languages. Real accuracy varies enormously between them, and the gap is widest exactly where Indian businesses need it.
The specific difficulty is not Hindi or English. It is the Hinglish people actually speak, where a sentence begins in one language, switches mid-clause, and carries English product names inside Hindi grammar. Many systems handle each language well in isolation and fall apart on the mixture, which is most real conversations.
There is no way to know how a provider performs on your callers except to test it on recordings of your callers. We do that during discovery, with two or three providers, before anything is chosen. It is unglamorous and it is the step that most determines whether the project works.
Where voice agents work, and where they do not
| Situation | Verdict |
|---|---|
| Missed inbound calls out of hours or during rushes | Strong, this is the clearest win |
| Appointment booking, rescheduling, reminders | Strong, structured and repetitive |
| Order status, delivery windows, opening hours | Strong, if the data is in a system the agent can read |
| Qualifying inbound leads before a human call | Good, and easy to measure |
| Complex complaints and anything emotional | Poor, route to a person quickly |
| Anything where a wrong action costs real money | Only with confirmation steps and limits |
| Low call volume already handled by a person | Do not build it |
What Crodo taught us
We build and run Crodo, a macOS voice assistant that takes dictation, answers questions about what is on the screen, writes meeting notes and acts in Gmail, Calendar and Slack. Running it taught us two things that shape every client build.
The first is that the interesting work is never the models. It is the plumbing: handling silence, deciding when the user has actually finished talking, recovering when transcription returns nonsense, and keeping the whole loop fast enough to feel like a conversation.
The second is about permissions and trust. A voice product asks for the microphone before it has demonstrated anything, and how you handle that moment decides your adoption rate. We wrote about that mistake honestly in four products in, what we got wrong.
Build or buy
Buy a platform when you want a standard receptionist live in days, your languages are well covered, and the per-minute price works at your volume. That is a genuinely good answer for a lot of businesses, and we will tell you when it is yours.
Build when the agent belongs inside your own product, when it must act in systems no platform integrates with, when your vocabulary or languages are poorly served, or when volume has made the platform margin bigger than a build. The AI integration page covers the case where voice is a feature of software you already run, and the clinic version of this decision, with the monthly arithmetic for each option, is in our guide to AI receptionists for clinics in India. For the general numbers, per minute and per build, with the four meters that run while the line is open, see our AI voice agent cost guide.
How a voice project runs
- Discovery, week 1. The calls that matter, recordings of real ones, the languages, the systems the agent must touch, and the rules for handing over to a person.
- Prototype, weeks 2 to 3. A working agent on a test number, scored on transcription accuracy, response latency and completion rate, using your recordings rather than a demo script.
- Build, weeks 4 to 10. Telephony or in-app integration, actions in your systems, monitoring, the evaluation set, and a rollout you can pause at any time.
Price is fixed after discovery. If the prototype shows the agent cannot handle your calls well enough, we stop there and say so, which is cheaper for both of us than finding out in month four.