Skip to content
AppWizards

04 / Voice AI & Conversational Interfaces

Voice AI development: agents that answer the phone, in the languages your customers speak.

AI call answering and voice agents for phone and WhatsApp, dictation inside your app, and call summaries, built by the team behind Crodo, our own voice assistant for Mac.

6 to 10 weeks · fixed price after a one-week discovery

Fig. 05 / How Crodo turns speech into actionsOUR PRODUCT
INMicrophone, streaming audio
PIPELINESpeech to text, language model, text to speech
OUTSpoken reply or an action in your apps
or a transfer to a person, with the context attached
Every stage streams, because the delay between them is what makes a bot feel like a bot

What we build

What you get: specific deliverables, not a list of capabilities.

01AI call answering and receptionistsAn agent that answers every call on the first ring, qualifies the caller, books the appointment and routes the rest, so nobody reaches a voicemail at 9pm or during a rush.
02Voice agents for phone and WhatsAppInbound and outbound calls and voice notes, in the languages your customers actually use, with a clean hand-off to a person and the context attached.
03Dictation and voice commandsFast, accurate speech to text inside your app, tuned to your vocabulary, with commands that do things rather than only transcribe them.
04Call and meeting summariesTranscripts, summaries, action items and CRM updates from calls and meetings, the way Crodo does it for meetings on a Mac.
05Latency and cost engineeringStreaming at every step, interruption handling, and the right model for each stage, so the conversation feels natural and the per-minute bill stays predictable.
06Monitoring and evaluationRecordings, transcripts and a scored test set of real calls, so you can prove the agent is still handling them properly a month after launch.

What a voice agent costs to run

Voice is priced by the minute, not by the token, and that single difference changes how you plan it. A well-built agent on Indian telephony costs roughly 2 to 4 rupees a minute to run, all-in. Here is where that goes.

StageIllustrative cost per minute
Speech to textabout $0.006
Language modelabout $0.007, varies with how much context you send
Text to speechabout $0.02 while speaking, so roughly half a call
Telephony30 to 70 paise on Indian numbers
All inroughly 2.5 rupees a minute

At 500 three-minute calls a month that is about 4,000 rupees of usage. At 2,000 calls, about 15,000. Prices move, so treat the method as the point rather than the figures, and note what the shape tells you: text to speech is usually the largest line, which is why choosing a cheaper voice for routine confirmations and a better one for the opening is a real cost lever rather than a detail.

Platform products commonly charge 8 to 30 rupees a minute. That is the same pipeline with a margin on it, and at low volume it is good value. The arithmetic reverses somewhere in the low thousands of calls a month.

How a voice pipeline works

Fig. 05 / How Crodo turns speech into actionsOUR PRODUCT
INMicrophone, streaming audio
PIPELINESpeech to text, language model, text to speech
OUTSpoken reply or an action in your apps

Three models in a row, each adding delay. If you run them one after another and wait for each to finish, the caller hears a pause long enough to say "hello?" into. Every stage has to stream: transcription emits words as they are spoken, the model starts generating before the caller finishes, and speech synthesis begins before the sentence is complete.

Latency is the whole product

Callers forgive a wrong answer more readily than a long silence. A wrong answer sounds like a person who is mistaken; a two-second gap sounds like a machine, and once someone decides they are talking to a machine, they stop cooperating with it.

So we budget latency by stage rather than measuring it at the end, and we build interruption handling in from the start. A caller who talks over the agent must be heard immediately, which means the agent has to stop speaking, discard what it was about to say, and listen. Retrofitting that is far harder than designing for it, because it touches every stage of the pipeline at once.

Indian languages, and the code-switching problem

Provider marketing lists dozens of supported languages. Real accuracy varies enormously between them, and the gap is widest exactly where Indian businesses need it.

The specific difficulty is not Hindi or English. It is the Hinglish people actually speak, where a sentence begins in one language, switches mid-clause, and carries English product names inside Hindi grammar. Many systems handle each language well in isolation and fall apart on the mixture, which is most real conversations.

There is no way to know how a provider performs on your callers except to test it on recordings of your callers. We do that during discovery, with two or three providers, before anything is chosen. It is unglamorous and it is the step that most determines whether the project works.

Where voice agents work, and where they do not

SituationVerdict
Missed inbound calls out of hours or during rushesStrong, this is the clearest win
Appointment booking, rescheduling, remindersStrong, structured and repetitive
Order status, delivery windows, opening hoursStrong, if the data is in a system the agent can read
Qualifying inbound leads before a human callGood, and easy to measure
Complex complaints and anything emotionalPoor, route to a person quickly
Anything where a wrong action costs real moneyOnly with confirmation steps and limits
Low call volume already handled by a personDo not build it

What Crodo taught us

We build and run Crodo, a macOS voice assistant that takes dictation, answers questions about what is on the screen, writes meeting notes and acts in Gmail, Calendar and Slack. Running it taught us two things that shape every client build.

The first is that the interesting work is never the models. It is the plumbing: handling silence, deciding when the user has actually finished talking, recovering when transcription returns nonsense, and keeping the whole loop fast enough to feel like a conversation.

The second is about permissions and trust. A voice product asks for the microphone before it has demonstrated anything, and how you handle that moment decides your adoption rate. We wrote about that mistake honestly in four products in, what we got wrong.

Build or buy

Buy a platform when you want a standard receptionist live in days, your languages are well covered, and the per-minute price works at your volume. That is a genuinely good answer for a lot of businesses, and we will tell you when it is yours.

Build when the agent belongs inside your own product, when it must act in systems no platform integrates with, when your vocabulary or languages are poorly served, or when volume has made the platform margin bigger than a build. The AI integration page covers the case where voice is a feature of software you already run, and the clinic version of this decision, with the monthly arithmetic for each option, is in our guide to AI receptionists for clinics in India. For the general numbers, per minute and per build, with the four meters that run while the line is open, see our AI voice agent cost guide.

How a voice project runs

  1. Discovery, week 1. The calls that matter, recordings of real ones, the languages, the systems the agent must touch, and the rules for handing over to a person.
  2. Prototype, weeks 2 to 3. A working agent on a test number, scored on transcription accuracy, response latency and completion rate, using your recordings rather than a demo script.
  3. Build, weeks 4 to 10. Telephony or in-app integration, actions in your systems, monitoring, the evaluation set, and a rollout you can pause at any time.

Price is fixed after discovery. If the prototype shows the agent cannot handle your calls well enough, we stop there and say so, which is cheaper for both of us than finding out in month four.

Frequently asked questions

Common questions about Voice AI & Conversational Interfaces.

How much does an AI voice agent cost per minute?

The usage cost of a well-built agent lands around 2 to 4 rupees a minute all-in on Indian telephony, covering speech to text, the language model, text to speech and the call itself. At 500 three-minute calls a month that is roughly 4,000 rupees of usage. The build is separate and quoted after discovery. Platform products often charge 8 to 30 rupees a minute, which is the same pipeline with a margin on top.

What is an AI receptionist, and is it worth it?

An agent that answers inbound calls, books appointments and routes what it cannot handle. It is worth it when you miss calls, which most clinics, salons, workshops and small service businesses do outside working hours and during rushes. It is not worth it when your call volume is low enough that a person already answers everything.

Which Indian languages can a voice agent handle?

English and Hindi are our daily work, and the speech providers we use cover most major Indian languages. The harder problem is code-switching, the Hinglish that people actually speak, where a sentence starts in Hindi and ends in English. We test accuracy on recordings of your own callers before choosing a provider, because provider quality varies far more by language than the marketing suggests.

How fast does it respond, and can you interrupt it?

Every stage streams, so a well-built agent starts replying in well under a second and can be cut off mid-sentence like a person. Latency is the single thing that decides whether callers accept it, which is why we measure and budget it stage by stage rather than treating it as a detail.

Can the voice agent transfer a call to a human?

Yes, and it should. It can transfer a live call or escalate a WhatsApp conversation with a summary and the caller's intent attached, so your team does not start from zero. Designing that hand-off well matters more than raising the share of calls the agent completes alone.

Should we build a voice agent or use a platform?

Use a platform if you want a standard receptionist live in days and the per-minute price is acceptable at your volume. Build when you need it inside your own product, when it must act in systems no platform integrates with, when your languages or vocabulary are not well served, or when volume makes the per-minute margin larger than the build.

What does it integrate with?

Phone numbers through the usual telephony providers, WhatsApp Business, and whatever the agent needs to act in: your calendar, CRM, order system or helpdesk. If a person has to open a system during the call today, that is the integration the agent needs.

How do you know it is working after launch?

With recordings, transcripts and a scored test set of real calls, run whenever the prompt, the model or the voice changes. Voice agents degrade quietly, and a call that ends politely without solving anything looks like a success in every metric except the one that matters.

Next step

Tell us what you want to build. We will tell you what it costs and how long it takes.

A free 30-minute call with an engineer, not a salesperson. You leave with a clear plan, a price range and an honest opinion on whether AI is the right tool for the job.