AI Agent Cost Per Task vs Your Price: Margin, Step Cap and Break-Even Volume

AI Agent Cost Per Task vs Your Price: Margin, Step Cap and Break-Even Volume

Work out what one AI agent task costs, what to charge for it, where to cap steps and the volume where your pricing breaks. Worked example inside.

Table of Contents

To see if an AI agent can make money, subtract its cost per task from the price you charge per task. That gap is your margin per task. It tells you how many steps a task can take before it loses money, and how many tasks a flat monthly plan can absorb before a customer costs more than they pay. In the worked example below, a 6-step invoice agent costs $0.144 per task at list prices and about $0.091 after prompt caching and model routing. At an assumed price of $0.25 per invoice, that leaves about $0.16 of margin per task before fixed costs.

Ready to build?

Fixed-price web and mobile MVPs from $3,460. Book a call or WhatsApp us.

The formula that decides your margin

An agent calls the model several times per task, and every call resends its instructions, its tool definitions and everything that has happened so far. So a task with more steps costs more than its share of the average, because the context keeps growing. Three formulas decide your margin.

  • Margin per task = price per task − (cost per task × (1 + retry rate))
  • Cost per task = sum of the cost of each step
  • Cost per step = (new input tokens × input price) + (cache writes × write price) + (cache reads × read price) + (output tokens × output price)

Cost per task: an invoice matching agent

The job: read a supplier invoice, find the matching purchase order, check the amounts and draft an entry for a human to approve. These assumptions are inputs for the calculation, not benchmarks.

  • 6 steps per task
  • 4,000 stable tokens per step (instructions and tool definitions)
  • 6,000 growing tokens per step on average (tool results and history)
  • 400 output tokens per step

Prices used in this example

Prices come from Anthropic's pricing page, read on 5 October 2026. Check them again on the day you set your price.

  • Claude Sonnet 5.5: $2 per million input tokens, $10 output, $2.50 for a 5 minute cache write, $0.20 for a cache read.
  • Claude Haiku 4.5: $1 input, $5 output, $1.25 cache write, $0.10 cache read.

At list price

Every step pays full input price on all 10,000 input tokens.

  • Input: 6 × 10,000 = 60,000 × $2/M = $0.120
  • Output: 6 × 400 = 2,400 × $10/M = $0.024
  • Per task: $0.144

With the stable prompt cached

The 4,000 stable tokens are written to the cache once, then read on the next 5 steps. Sonnet 5.5 accepts prompts from 512 tokens for caching, so this works.

  • Cache write: 4,000 × $2.50/M = $0.010
  • Cache reads: 5 × 4,000 = 20,000 × $0.20/M = $0.004
  • New input: 6 × 6,000 = 36,000 × $2/M = $0.072
  • Output: 2,400 × $10/M = $0.024
  • Per task: $0.110

With easy steps routed to Haiku 4.5

Four of the six steps are simple: choose the next tool, read a result. They go to Haiku 4.5. The two hard steps, matching amounts and drafting the entry, stay on Sonnet 5.5. One trap here. Anthropic's prompt caching docs set the minimum cacheable prompt length at 4,096 tokens for Haiku 4.5. Our 4,000 stable tokens fall under it, so Haiku pays full input price on every step. Only route a step if your evals show the smaller model gets it right. A cheap step that picks the wrong tool adds steps, and extra steps eat the margin you are trying to protect.

  • Haiku 4.5 (4 steps, no cache): no cache write, no cache reads. Input: 40,000 × $1/M = $0.0400. Output: 1,600 × $5/M = $0.0080. Subtotal: $0.0480.
  • Sonnet 5.5 (2 steps, cached): cache write: 4,000 × $2.50/M = $0.0100. Cache reads: 4,000 × $0.20/M = $0.0008. Input: 12,000 × $2/M = $0.0240. Output: 800 × $10/M = $0.0080. Subtotal: $0.0428.
  • Per task: $0.0908, about $0.091.
  • If your real prompt is close to 4,096 tokens, count it with the provider's token counter. Once it passes the minimum, Haiku caches it too.

Margin at four volumes

Assumed price: $0.25 per invoice processed. This is a figure for the calculation, not a market price or a recommendation. The figures are model costs only, before retries and fixed costs. With per-task pricing, the margin rate stays the same at every volume: about 64% optimised, about 42% at list price. Volume grows the dollars, but it never fixes a thin margin per task. Two things do break the model: tasks that take too many steps, and flat plans with no task limit.

  • 1,000 tasks per month. List: $144. Cached: $110. Cached + routed: $91. Margin at $0.25/task (cached + routed): $159.
  • 10,000 tasks per month. List: $1,440. Cached: $1,100. Cached + routed: $910. Margin at $0.25/task (cached + routed): $1,590.
  • 50,000 tasks per month. List: $7,200. Cached: $5,500. Cached + routed: $4,550. Margin at $0.25/task (cached + routed): $7,950.
  • 100,000 tasks per month. List: $14,400. Cached: $11,000. Cached + routed: $9,100. Margin at $0.25/task (cached + routed): $15,900.

The step cap: when one task starts losing money

Each extra Sonnet step in this example costs at least about $0.017 (breakdown below). The $0.159 margin covers about 9 extra steps at that rate. So a task that loops to around 15 steps loses money. In practice it breaks earlier, because context grows with each step and later steps cost more than the average. The rule: max steps = planned steps + (margin per task ÷ cost of one extra step), then set your cap well below that. When the agent hits the cap, it stops and hands the task to a human. A capped task costs you a review. An uncapped loop costs you money on every run.

  • Cache read: 4,000 × $0.20/M = $0.0008
  • New input: 6,000 × $2/M = $0.012
  • Output: 400 × $10/M = $0.004
  • Total: about $0.017

Want us to build it?

Fixed-price web and mobile MVPs from $3,460. Book a call or WhatsApp us.

The volume where a flat plan breaks

Say you sell a plan at $49 per customer per month with unlimited invoices (again, an assumption for the math). Above the break-even line, each extra invoice from that customer costs you money. Your heaviest users are often your best customers, and they are the ones who cross it first. The fix is to include a set number of tasks below the break-even line, then charge per task above it. Or price per task from day one, as in the volume figures above.

  • Break-even per customer = $49 ÷ $0.091 = about 540 invoices a month
  • At list price: $49 ÷ $0.144 = about 340 invoices a month

Costs outside the model bill

Retries, hosting, a background queue, tracing, a vector database and third-party APIs come on top, each priced on the provider's own page. Your minimum monthly volume = fixed monthly costs ÷ margin per task: at $0.159 per task, every $100 of fixed costs needs about 630 tasks to cover it.

Four ways to widen the margin

Each of these lowers your cost per task without touching your price.

  • Cap the steps. Use the rule above. It protects margin from the first day.
  • Cache above the minimum. Check each model's minimum cacheable length before you count on the discount.
  • Trim the context. Summarise old tool results instead of resending them raw. This lowers the cost of every late step.
  • Batch what can wait. The Batch API from Anthropic and OpenAI takes 50% off input and output tokens and stacks with caching. Results can take up to 24 hours, so it suits evals, not live loops. 50 test tasks per release and 8 releases a month make 400 tasks × $0.091 = about $36.40, or about $18.20 batched.

Get your real numbers first

These numbers come from assumptions. Yours come from 10 real tasks run through a prototype, with the token counts read from each API response. If you already have an agent prototype built with an AI tool, read our guides on taking an AI prototype to production. On a free call, we look at your idea and give you a fixed quote. Book a call.

How do I price an AI agent per task?

Measure the cost per task on real runs and multiply it by 1 plus your retry rate. Then set a price that leaves enough margin to cover your fixed costs at your expected volume. Check that the margin still holds when a task takes several extra steps.

How many steps should an AI agent be allowed per task?

Take your planned steps and add the margin per task divided by the cost of one extra step. Set your hard cap well below that number and send capped tasks to a human.

Does prompt caching work on every model?

No. Each model has a minimum cacheable prompt length. On Anthropic's API it is 512 tokens for Sonnet 5.5 and 4,096 tokens for Haiku 4.5. Below the minimum, you pay full input price.

Why does an agent cost more per task than a chatbot?

An agent calls the model several times per task and resends the growing context at each step. A chatbot usually makes one call per reply.

Is a flat monthly plan safe for an AI agent product?

Only with a task limit. Divide the plan price by your cost per task: that is the number of tasks a customer can run before they cost you more than they pay.

Ready to build yours?

Fixed-price web and mobile MVPs from $3,460. Book a call or WhatsApp us.

Related Articles

Ready to ship your MVP?

Fixed-price builds from $3,460 · Post-launch support from $500/mo

Frequently Asked Questions

Who Is Behind BuildMVPFast?

BuildMVPFast is led by a passionate team of Creators, Designers, and Developers. Together, we deliver innovative software solutions tailored to each client's unique needs. Our vision is to help clients unlock their full potential by providing a rapidly built, high-quality MVP that serves as the perfect launchpad for their business idea.

Who Owns The Code, Design And Intellectual Property?

How Long Does It Take To Build An MVP?

What Is Your Development And Delivery Process?

Do You Offer Post-Launch Support And Maintenance?