Back to Journal
AI Strategy 11 min read

How Much Does It Cost to Run an AI Agent?

The unit economics buyers ask about after the build quote — cost per task, why agentic loops multiply token spend, what caching and routing actually save, and the monthly math at 100, 10,000, and 1M tasks.

Key Takeaways

  • A single chatbot turn costs well under a cent in model spend. A multi-step agent task that plans, calls tools, and retries typically costs $0.10-$0.50 — roughly 20-100x more, for the same underlying model.
  • Agentic loops are expensive because cost grows faster than step count: every step re-sends everything the earlier steps produced. An 8-step loop costs about 37x one chatbot turn, not 8x.
  • Prompt caching is the single highest-leverage lever. Anthropic bills a cache hit at 0.1x the base input price, and on a realistic agent loop that roughly halves the task cost with no quality change.
  • Caching plus model routing — small models for the mechanical steps, a frontier model only for the hard ones — routinely takes a $0.22 task under $0.08, a 60%+ cut before you touch the product.
  • Human review usually costs more than tokens at mid volume. At 10,000 tasks a month with a 12% review rate, reviewers cost roughly 4x the model bill, which is why eval work that lowers the review rate has the best ROI in the whole system.
  • Below roughly 10,000-25,000 tasks a month, run cost sits inside the usual 20-40% of build value per year. Above that, usage dominates and you should be budgeting per task, not as a percentage of the build.

A completed AI agent task – one where the agent plans, calls a few tools, and returns a verified answer – typically costs $0.10 to $0.50 in model spend before optimization, and $0.03 to $0.15 after prompt caching and model routing. A single chatbot turn on the same model costs well under a cent. That gap – roughly 20x to 100x for the same underlying model – is the number that surprises buyers who signed off on a build quote and then opened their first full month of invoices.

This is the run-cost question, not the build-cost question. If you are still scoping the one-time number, our AI total cost of ownership guide covers the full three-year picture. This article goes one level deeper into the line that scales with success: what one agent task costs, why agentic loops multiply token spend, and what the monthly bill looks like at 100, 10,000, and 1,000,000 tasks.

What does one AI agent task actually cost?

"Cost per task" is only meaningful once you say what a task is. The five workload shapes below all use large language models, but they differ by two orders of magnitude in price. Figures are 2026 model spend at mid-tier pricing (roughly $3 per million input tokens and $15 per million output tokens), excluding human review and fixed infrastructure.

Workload unitTypical tokens per unitModel cost per unitWhat drives it
Simple chatbot turn~800 in / ~250 out$0.003–$0.01System prompt length, history depth
RAG query (retrieval-grounded answer)~6k–15k in / ~500 out$0.02–$0.06Chunks retrieved, chunk size, reranking
Multi-step agent task (5–10 calls, tools)40k–120k cumulative$0.15–$0.50Steps per task, context re-sending, retries
Long-running research or coding agent300k–1M+ cumulative$1.50–$10+Loop depth, volume of tool output
Voice agent minute~4–6 turns per minute$0.05–$0.20 / minSpeech-to-text and text-to-speech vendors, not the LLM

Two things in that table matter more than the absolute numbers. First, the multi-step agent row is the one most people mean when they say "AI agent," and it is 20–50x the chatbot row. Second, the voice row is the only one where the language model is not the dominant cost – transcription and synthesis usually are – so voice budgets that only model tokens are wrong by a factor of three.

Why do agentic loops multiply token spend?

Because an agent re-reads its own history on every step. Each call in the loop re-sends the system prompt, the tool schemas, and every tool result produced so far. Input tokens therefore accumulate quadratically in the number of steps, not linearly.

Work it through. Say a support-resolution agent starts with a 2,000-token prefix (system prompt plus tool definitions) and each step adds about 1,500 tokens of tool output and reasoning to the transcript. Step one sends 2,000 input tokens; step eight sends 2,000 plus seven steps' worth of accumulated context, about 12,500. Summed across eight steps that is roughly 58,000 input tokens, plus about 3,200 output tokens at 400 per step. At $3 and $15 per million:

  • Input: 58,000 x $3 / 1,000,000 = $0.174
  • Output: 3,200 x $15 / 1,000,000 = $0.048
  • Total: $0.222 per task

Eight independent chatbot turns would have cost about $0.048 in total. The loop cost roughly 37x one turn, not 8x, because step eight pays for everything steps one through seven produced. Then add the two multipliers nobody models: retries and escalation. If 15% of tasks fail a validation check and re-run, your effective cost is $0.222 x 1.15 = about $0.26. If 5% escalate to a frontier model, add a few cents more. A quoted "$0.20 per task" that ignores retries is understated by 20–30% before it reaches production.

This is also why context-window size is a cost decision, not just an engineering one. Long context is priced per token like everything else, so an agent that stuffs a whole document into the prefix pays for that document on every step of every task. Retrieving three relevant chunks instead of pasting forty pages is often a 10x cost change with no quality loss – the tradeoff we unpack in how to build a RAG system.

How much do caching and model routing actually save?

The good news is that the exact thing making agents expensive – a repeated, stable prefix – is what prompt caching was built for. According to Anthropic's published API pricing, 2026, a cache read is billed at 0.1x the base input price – 10% of what the same tokens cost uncached – with a one-time write premium of 1.25x for the five-minute cache.

Apply that to the loop above. Suppose 45,000 of the 58,000 input tokens are cache hits on a stable prefix, 13,000 are genuinely new, and about 3,500 are written once:

  • Cache reads: 45,000 x $0.30 / 1,000,000 = $0.014
  • Uncached input: 13,000 x $3 / 1,000,000 = $0.039
  • Cache writes: 3,500 x $3.75 / 1,000,000 = $0.013
  • Output (unchanged): $0.048

New total: about $0.114 per task – caching alone cut the bill roughly in half, with no change to the output. The engineering condition is that your prefix must actually be stable, so put timestamps, user IDs, and anything else volatile at the end of the prompt, not the beginning.

Model routing is the second lever. Most steps in an agent loop are mechanical – classify an intent, extract fields, format a tool call, summarize a result – and run fine on a small model priced around $1 per million input tokens rather than $3 or $5. Route 60–70% of tokens to the cheap tier and keep the planning and final-answer steps on the capable model, and the blended input price drops by about half again. Combined, caching and routing routinely take our $0.22 task under $0.08. Both moves need an evaluation set to prove quality held; the full engineering playbook is in how to reduce LLM API costs.

What does human review add per task?

For most business agents, humans cost more than tokens. Any workflow with real consequences keeps a reviewer on the cases the agent should not decide alone, and that review is a recurring labor line priced at volume x review rate x minutes x loaded hourly rate.

At a 12% review rate, eight minutes per review, and a $35 fully loaded hourly rate, each reviewed task costs about $4.67 – roughly 40x the model cost of the task itself. Spread across all tasks, that is $0.56 per task in review labor against $0.11 in tokens. The implication is blunt: optimizing prompts to shave a cent off inference is nearly worthless if a two-point reduction in review rate is available instead. Evaluation and reliability work is the highest-ROI cost lever in the system, not a quality nicety – which is the argument behind evaluating and testing AI agents properly before launch.

What does it cost at 100, 10,000, and 1M tasks a month?

Every agent also carries a fixed floor: hosting, the vector database, observability and trace storage, scheduled evaluation runs, and someone maintaining prompts as providers ship model updates. Call it $2,000–$8,000 a month for a single production agent. Now scale the variable parts at $0.13 per task and a 12% review rate:

Monthly volumeModel spendHuman reviewFixed floorTotal per month
100 tasks$13~$56$2k–$8k$2k–$8k
10,000 tasks$1,300~$5,600$2k–$8k$9k–$15k
1,000,000 tasks$130,000Not viable at 12%$8k–$25k$140k–$250k+

Read the rows, not the totals. At 100 tasks a month, tokens are a rounding error and you are paying almost entirely for the privilege of having a maintained system at all – which is why pilots that never scale look so bad on a cost-per-task basis. At 10,000, review labor is about 4x the model bill and the whole thing still fits comfortably inside the usual planning anchor of 20–40% of build value per year: on a $100k build, that band is $20k–$40k a year, and this agent runs $108k–$180k. Note what that means – a single high-volume agent can consume the entire run-cost budget you allocated for the product.

At 1M tasks a month, two things break. A 12% review rate would mean 120,000 reviews and roughly $560,000 a month, so confidence-based routing has to push the review rate toward 1% or the business case collapses. And a 30% inference optimization is worth about $470,000 a year – enough to fund a dedicated engineer several times over. Percentage-of-build budgeting stops being useful here. Above roughly 10,000–25,000 tasks a month, price the agent per task and put a margin on it, the way you would any other unit-cost business.

How should you model agent run cost before you build?

Estimate bottom-up in four terms, in this order: steps per task, tokens per step, tasks per month, and review rate. Multiply the first two against published per-token prices for a raw cost per task, apply a 40–50% caching discount and a routing discount only if your architecture genuinely supports them, then multiply by volume and add the fixed floor and review labor. Finally, run the same model at 10x your expected volume – that is the number that tells you whether the product survives its own success.

The three assumptions worth arguing about in that model are steps per task (agent designs vary from three to thirty, and the cost difference is quadratic), retry rate (rarely modeled, usually 10–20%), and review rate (the single biggest line at mid volume). Ask any vendor quoting you an agent for all three numbers explicitly. A partner who has actually run an agent in production will have them; one who has only demoed will quote a per-task figure with no steps behind it.

That is how we scope agent work at Game Changer Labs: the per-task math ships with the proposal, with steps, retries, and review rate named as assumptions you can challenge. You can see how we structure that on our services page – and if the quote in front of you right now has a build price but no cost per task, you are holding half a number.

Frequently Asked Questions

How much does it cost to run an AI agent per task?

For a typical business agent in 2026, budget $0.10 to $0.50 per completed task in model spend before optimization, and $0.03 to $0.15 after prompt caching and model routing. The wide range comes from how many steps the agent takes: a single-turn chatbot reply costs well under a cent, a retrieval-grounded answer costs one to five cents, and a multi-step agent that plans, calls three or four tools, and occasionally retries costs ten to fifty times a single turn. Tokens are only part of the bill — add human review, observability, and evaluation to get the real number.

Why do AI agents cost so much more than chatbots?

Because an agent re-reads its own history on every step. Each model call in the loop re-sends the system prompt, the tool schemas, and every tool result produced so far, so input tokens accumulate quadratically rather than linearly. An eight-step agent loop consumes roughly 58,000 input tokens where eight independent chatbot turns would consume about 6,400. Add tool-call overhead, failed steps that get retried, and larger context windows, and a single agent task can cost 20 to 100 times one chatbot turn on exactly the same model.

How much does prompt caching save on agent costs?

Roughly 40-50% of the total cost of a typical agent task, and more for agents with large stable system prompts or long retrieved documents. Anthropic's published pricing bills a cache read at 0.1x the base input token price — 10% of the uncached rate — with a one-time write premium of 1.25x for the five-minute cache. Because an agent loop re-sends the same prefix on every step, the prefix is exactly the content caching was designed for. The catch is that caching only pays off if your prompt prefix is genuinely stable, so put the volatile parts at the end.

What does it cost to run an AI agent at 10,000 tasks a month?

At a realistic post-caching cost of about $0.13 per task, 10,000 tasks a month is roughly $1,300 in model spend, or $15,600 a year. That is usually the smaller half of the bill. Add a fixed operating floor of about $2,000-$8,000 a month for hosting, observability, evaluation runs, and prompt maintenance, plus human review: at a 12% review rate and eight minutes per review, reviewers cost around $5,600 a month. Total run cost lands near $9,000-$15,000 a month at that volume.

Is it cheaper to use a small model or a big model for agents?

Neither, exclusively — the cheapest production pattern is routing. Most steps in an agent loop are mechanical: classify an intent, extract fields, format a tool call, summarize a result. Those run fine on a small model priced around $1 per million input tokens, roughly a fifth of a frontier model's rate. Reserve the expensive model for the planning and final-answer steps where reasoning quality actually changes the outcome. Routing 60-70% of tokens to the cheap tier typically cuts the blended token price by about half without a measurable quality drop, provided you have an evaluation set to prove it.

How do I estimate AI agent running costs before I build?

Model it bottom-up in four terms: steps per task, tokens per step, tasks per month, and human-review rate. Multiply the first two by published per-token prices to get a raw cost per task, apply a 40-50% caching discount and a routing discount if your architecture supports them, then multiply by monthly volume. Add a fixed floor for hosting, monitoring, and evaluation, and add review labor as volume times review rate times minutes times loaded hourly rate. Sanity-check the result against the usual planning anchor of 20-40% of build value per year — if you are far above it, usage is your dominant cost and you should be pricing per task instead.

Free Tools

Game Changer Labs

Tell us what you're building — book a free scoping call.

Pick a time that works and walk us through your project — 30 minutes, straight to the point. You leave with a concrete plan, timeline, and cost. No sales pitch — if we're not the right fit, we'll say so.

Keep Reading

Get new playbooks by email

Occasional, no-fluff field notes on building production AI — new guides and tools, straight to your inbox. Unsubscribe anytime.

Published: August 7, 2026Game Changer Labs