Skip to content
AppWizards

Costs · AI agents

What an AI Agent Costs to Run: A Worked Example

An AI support agent on 5,000 tickets a month costs roughly $190 to $320 to run. The full arithmetic, the costs nobody budgets for, and five ways to cut it.

Yash Rai · 24 August 2026 · updated 14 September 2026 · 11 min read

Racks of servers in a dark data centre, threaded with amber and cyan fibre optic cables and rows of green status lights.

An AI support agent handling 5,000 tickets a month costs roughly $190 to $320 a month to run, of which about $117 is model usage and the rest is hosting, search and monitoring. That works out at just over 2 cents per ticket.

Here is the full arithmetic, and then the more useful part: why the token cost is the number everyone models and the one that matters least.

The two costs, and why people conflate them

Building an agent is a one-time cost. Running it is a monthly bill that grows with every request, which is what makes AI different from ordinary software. A feature you ship once and forget has a marginal cost of roughly zero. An agent has a meter on it.

Most cost estimates you will find online model the meter carefully and stop there. That is the wrong half to obsess over, for reasons the numbers below make obvious.

The worked example: a support agent

Assume a helpdesk receiving 5,000 tickets a month, and an agent that reads each ticket, searches your help articles, and drafts a reply.

Each ticket involves roughly:

  • Reading and context. The ticket, the customer's history and the relevant help articles come to around 4,000 tokens of input. A token is roughly three quarters of a word, and models charge per million tokens.
  • The reply. Around 400 tokens of output. Output tokens usually cost several times more than input tokens.
  • Retries and checks. Some requests run twice: a quality check, a retry after a timeout, a second search when the first found nothing. A realistic overhead is 30 percent.

At an illustrative mid-tier price of $3 per million input tokens and $15 per million output tokens:

ItemMonthly amount
Input: 5,000 tickets x 4,000 tokens x 1.3 overhead26 million tokens, about $78
Output: 5,000 tickets x 400 tokens x 1.3 overhead2.6 million tokens, about $39
Search index and hosting for the agent service$50 to $150
Monitoring and logging$20 to $50
Totalroughly $190 to $320 a month

Model prices move often, so treat the method as the point rather than the exact figures. What does not move is the shape.

Cost per ticket, and what happens at other volumes

Divide it out and each ticket costs about 2.3 cents in model usage. That single number is more useful than the monthly total, because it scales and the monthly total does not.

Tickets per monthModel usageFixed costsTotal
500about $12$70 to $200$82 to $212
5,000about $117$70 to $200$187 to $317
50,000about $1,170$150 to $400$1,320 to $1,570

Look at the 500-ticket row. The model usage is $12 and the bill is up to $212, because the fixed costs dominate completely. At 50,000 tickets the position reverses and usage is most of the bill.

That inversion is the whole planning insight. At low volume you are paying for infrastructure, not intelligence. An agent is not expensive to run at small scale; it is expensive to justify at small scale, which is a different problem with a different answer.

What a ticket costs three ways

The number that decides the project is not the agent's cost in isolation. It is the agent's cost against the alternatives you already run.

How the ticket gets handledIllustrative cost per ticketWhat you also get
A person answers it$0.60 to $1.20Judgement, and an answer for anything
A help centre article answers itNear zeroOnly for questions someone wrote up
The agent answers itAbout $0.02 in model usageEverything the help centre covers, phrased for the actual question

The per-person figure assumes a support salary of roughly 25,000 to 40,000 rupees a month and around 400 tickets handled, which is a common shape in Indian support teams. Adjust it to yours; the ratio is what matters and the ratio is not close.

That gap is why agent projects look irresistible on a slide and disappoint in practice. The model cost genuinely is trivial. The project still fails if the agent resolves 30 percent of tickets rather than 60, or if the engineering to keep it at 60 costs more than the people it replaced.

A rising line chart on a dark gridded screen, with volume bars along the bottom.
The line most teams watch. The one that decides the project is usually flat and sits somewhere else.Photo: Unsplash

The four costs nobody puts in the spreadsheet

Every estimate models tokens. Almost none of them model these, and these are where projects actually go wrong.

Engineering time after launch. An agent is not finished when it ships. Prompts get tuned, retrieval gets fixed, edge cases get handled, and a model version gets deprecated with three months' notice. Budget a few days a month of a developer's attention, indefinitely. At Indian rates that is comfortably larger than the model bill in our example, and at Western rates it dwarfs it.

The humans still in the loop. The 20 to 40 percent of cases the agent hands to a person still cost people time. Worse, they are the hard cases, so they take longer per ticket than the average did before. Count the agent's value on the cases it finishes, not the cases it touches.

Evaluation runs. Every prompt change, model upgrade or data refresh should be scored against a test set before it ships. If your set has 200 cases and you run it twenty times a month, that is 4,000 extra model calls. Small in money, easy to forget, and skipping it is how quality quietly degrades.

Peak capacity. Model providers cap requests per minute. A campaign that triples your ticket volume for a week needs queueing and retry logic, or the agent falls over at exactly the moment it was supposed to prove itself.

A numbered network patch panel, with blue and grey cables plugged into labelled ports.
Everything on an agent's bill is metered per request. Almost nothing about the work around it is.Photo: Unsplash

Where the number actually moves

If you want to change the bill, these are the levers in order of how much they matter.

Fig. A / What moves an agent's running costIn order of effect
BIGGESTHow much context you send
LARGEWhich model does which step
SMALLPrompt wording and output length

Context size dominates because input tokens are the bulk of the volume. Halving what you send by improving retrieval halves most of the bill. Model choice comes second and is nearly as powerful when the work is split properly. Prompt micro-optimisation comes a distant third, and it is where most teams spend their time.

Five ways to cut the bill, and what each is worth

  1. Use a smaller model for the easy steps. Classifying a ticket does not need the model that writes the reply. Splitting the work typically cuts total model cost by a third to a half, and it is the single highest-return change available.
  2. Cache what repeats. In support, the same twenty questions arrive constantly. Serving a checked answer from cache costs nearly nothing, and in a mature support agent this can absorb a large share of traffic.
  3. Send less context. Better search over your help articles means fewer tokens per request. Retrieval quality is a cost lever, not only a quality lever, and it improves answers at the same time.
  4. Batch the background work. Anything that does not need an instant answer can run in cheaper batch modes overnight. Summarising yesterday's tickets does not need to happen in real time.
  5. Measure cost per resolved case, not per request. A change that cuts cost per request by 20 percent but drops resolution from 60 to 45 percent has made your economics worse. The metric that matters is the cost of a case the agent finished correctly.

What running our own products taught us

We build and run AI products of our own, which is a different education from building them for clients and handing them over.

The lesson that transfers most directly comes from Decornoa, which generates interior redesigns from a photograph of a room. The hard problems there were never the model. They were generation cost, queueing under load, and quality control, which is to say: the plumbing around the model, not the model itself. Every serious cost surprise we have seen since has been in that same category.

The second lesson comes from operating Crodo, a macOS voice assistant, where usage is continuous rather than bursty. When a product is used all day, small per-request inefficiencies that look like rounding errors in a spreadsheet compound into the largest line on the bill.

Neither of those is something you learn from a pricing page. It is why our estimates start from a prototype running on real data rather than from arithmetic alone.

When the answer is do not build it

At 200 tickets a month, do not build this. The fixed costs alone exceed what the agent could possibly save, and the honest recommendation is a better help centre, a few canned replies and a well-designed contact form.

Roughly speaking, an agent starts making sense somewhere above a thousand tickets a month, and becomes obvious above five thousand. Below that the maths says no, and we will tell you so on the call rather than after the invoice.

The same applies to any process where a person could not check the output. If nobody can tell whether an answer was right, cost per resolved case is not measurable, and you are buying activity rather than outcomes.

Questions people ask

Is it cheaper to run an open model ourselves?

Sometimes, at high volume. Self-hosting trades a per-request bill for a fixed GPU bill plus engineering time. It usually starts winning well above the volumes in this article, and it stops being about money and starts being about data residency and control long before that.

How much does an AI agent cost to build, as opposed to run?

Most agent projects take four to eight weeks. The build cost is a one-time project figure, quoted after a one-week discovery; across the Indian market an agent that only answers costs 3 to 8 lakh and one that acts in your systems 8 to 25 lakh. The ranges, and a worked estimate for this same support agent, are in what it costs to build an AI agent. The running cost is what this article models, and the two are worth keeping separate in any budget you present.

Do prices only go down?

Per-token prices have fallen consistently, but bills often have not, because teams spend the savings on larger context and more capable models. Plan for your usage to grow into the price cuts rather than assuming next year is cheaper.

Does a longer context window make agents cheaper?

No, it makes them more expensive if you use it. A larger window is permission to send more context, and you are billed for what you send. The teams whose bills surprise them are usually the ones who filled a bigger window because it was there rather than because the extra context improved answers.

How do we budget for an agent when model prices keep changing?

Budget on cost per resolved case rather than on a monthly total, and re-check it quarterly. A per-case figure survives a price change, a model swap and a volume change, and it is the only number that stays comparable across all three.

What about caching offered by the model provider?

Provider-side caching of repeated context can meaningfully reduce input costs when your prompts share a large fixed prefix, which is common in support agents. It is worth designing for from the start, since retrofitting it means restructuring prompts.

How to instrument this yourself

If you already run an agent and do not know what it costs per resolved case, that is the first thing to fix, and it is a day of work rather than a project.

Log four fields on every single request: input tokens, output tokens, which model handled it, and the outcome. Outcome is the one people skip and the only one that makes the rest meaningful, because cost per request is a vanity metric and cost per resolved case is a business one.

With those four fields you can answer the questions that actually change decisions. Which ticket types cost the most and resolve the least. Whether your retries are a rounding error or a third of the bill. Whether last month's prompt change paid for itself. Whether a cheaper model on the classification step would cost you anything in quality.

Almost every agent we are asked to review is missing the outcome field, which means nobody in the building can say whether the thing is working.

How we estimate this before you commit

Every agent project we take on starts with a prototype on your real data, scored against a test set of real examples. That prototype produces two numbers: how often the agent is right, and what it costs when it is. The decision to go to production is made with both on the table.

We would rather show you a small number that says no in week two than a large invoice that says yes in month six.

Where to go next

If you are working out whether an agent is the right shape of solution at all, that is what an AI readiness audit answers in two weeks. If you already know and want to know how we build them, see AI agent development. If the agent belongs inside software you already run, that is AI integration.

For the build-side budget rather than the running cost, our app development cost guide for India covers what moves a project quote. If the agent lives on WhatsApp, Meta's per-message fees are a third bill on top of the two here, and the arithmetic for it is in our WhatsApp AI agent cost guide.

Next step

Tell us what you want to build. We will tell you what it costs and how long it takes.

A free 30-minute call with an engineer, not a salesperson. You leave with a clear plan, a price range and an honest opinion on whether AI is the right tool for the job.