How to measure the cost per operation of an AI feature

A feature that costs pennies in the demo costs a salary at a thousand users. How to measure the cost per operation, and how to put a hard ceiling on it.

By Francisco José Fernández-Medina LópezPublished on 8 min
519 48

An AI feature that costs pennies in the demo costs a salary at a thousand users. The problem is not that nobody knows this: it is that the cost does not surface until the feature is in production, and by then the product's pricing is already published.

This is what we learned putting a price on the content planner inside PlanVortex, which writes posts and generates images, and it is what we ask before budgeting anything similar.

The unit is not the token, it is the operation

Nobody buys tokens. The customer buys "a week of posts", "this document classified" or "this invoice extracted". That is the unit to measure, and it is almost never a single model call.

A real operation includes:

  • the visible call (the one generating the text or the image),
  • the intermediate passes the user never sees (the one deciding what to generate, the one validating the result),
  • the input, usually the fastest growing part and the least measured,
  • and the retries, which are part of the cost even when they are the provider's fault.

Measuring per call gives you a prettier number with no use. Measuring per operation gives you the number you can build a price on.

Cost is not set by the model, it is set by the source

This is the finding that changed our design the most, and it has nothing to do with which model you use.

In our planner, the same request ("build me a week of posts") costs wildly different amounts depending on where the material comes from:

  • If the AI generates the images, a seven-post plan lands around 519 credits.
  • If the customer supplies the photos (their own, or their shop's catalogue), the same plan costs 48.

Ten times less, and the difference is neither the model nor the prompt: it is a decision the user makes in the first step of the wizard. Input weighs too, and that is not obvious either: a long article as a source adds several thousand input tokens, and each image entering a multimodal pass costs roughly what a generated paragraph does.

The practical consequence is that a flat per-operation price lies the moment the source changes. Either the pricing distinguishes the source, or you are subsidising the expensive case with the cheap one until a customer shows up who only uses the expensive one.

Estimating and charging are different things (and they have to agree)

There are two pieces of arithmetic here, and it pays to be clear about them from the start:

  • The estimate, before running, to show the user what they are about to spend and to check their balance covers it. It is necessarily approximate: you do not yet know how many tokens the model will return.
  • The charge, afterwards, using the real cost the provider reports for that call.

Both have to come from the same tariff table. When they do not (when the estimate lives in the dashboard and the charge lives on the server) they drift apart on the first change, and an estimate that does not match the charge is worse than no estimate at all: the customer reads it as a quote and the invoice contradicts it.

We solve this by duplicating the table on purpose on both sides, with a test on each side that fails if they stop saying the same thing. It is not elegant; it is what prevents the expensive failure.

The unit the customer sees cannot be the dollar

Model prices change several times a year, almost always downwards, and your product's pricing cannot change every time. So consumption is translated into your own unit (credits, in our case) with a conversion rate maintained internally.

Two effects, both good:

  1. The customer sees a stable figure and can compare months. A text is 2 credits and an image is 70, which tells them at a glance where their consumption goes.
  2. You can change provider without touching the customer-facing price, which is exactly what you cannot do once you have published a price in dollars per million tokens.

And one warning, because this is where a lot of products' numbers break: if the cost is in dollars and the price is in euros, the exchange rate is part of your margin. That is not an accounting detail, it is a variable that moves on its own.

The per-account cap is not a safety net, it is what makes the price possible

Without a cap, a customer's worst possible month has no limit: a bad loop, an integration retrying without brakes, or simply a customer using the feature far more than expected. And you find out from the provider's invoice, which arrives when nothing can be done.

With a cap per account and period, checked before each call rather than after, the worst possible month is a specific number. That is what lets you write a price down and stand behind it.

It is also worth leaving an exit for the customer who legitimately outgrows the cap: let them bring their own provider key and pay the provider directly. They stop consuming your allowance, they stop being a margin problem, and it is still your product.

Changing model is not changing one line

Every few months a cheaper or better model appears, and the temptation is to change the identifier and deploy. The effect on the invoice is visible the next day; the effect on quality is not visible until a customer complains, and by then nobody knows whether it got worse because of the model or because of something else.

What makes it decidable is having your own set of cases: twenty or thirty real examples, with what counts as a good answer, kept from the beginning. With that, switching models is a half-hour comparison. Without it, it is a bet you pay for in support.

And the part nobody budgets: human review

What a model produces enters as a draft. Publishing it, sending it or signing it off is still a person's decision, and that belongs in the cost of the feature: somebody reviews, and that somebody takes time.

It is also why there are tasks where automating does not pay off yet: if reviewing the output costs more than doing the work by hand, the feature is an expense dressed up as a saving. It is the first thing we look at when someone asks us for one, and when that is the answer, we say so in the first conversation.

What to ask before switching an AI feature on

  1. What is the operation the customer is buying, and what does the whole thing cost?
  2. How much does that cost change with whatever the user brings?
  3. Do the estimate they see and the amount they are charged come from the same place?
  4. What is the worst possible month for a single customer?
  5. What do you compare against the day you change model?

All five are answered before writing the feature, not after. That is what our applied AI service includes: the specific feature, measured per operation and with a ceiling, not "AI" in the abstract.

Frequently asked questions

How do you calculate the cost of an AI feature?
Per complete product operation, not per model call. An operation includes the input (which grows with whatever the user brings), the output, the retries and the intermediate passes the user never sees. Measuring per call gives you a prettier number with no use: nobody buys calls, they buy 'a week of posts' or 'this document classified'.
Why charge in credits instead of euros or tokens?
Because model prices change every few months and your pricing cannot change with them. Your own unit lets you move the conversion rate internally without touching the customer-facing price, and lets the customer see consumption in something they understand. Nobody understands tokens: a token is not a business unit.
What is a per-account cap and why is it needed?
A consumption limit per customer and period, checked before each call. Without one, a customer with a bad loop, or with legitimate but enormous usage, eats a whole month's margin, and you find out from the provider's invoice, which arrives when nothing can be done about it. With a cap, the worst possible month is a number you can put in a budget.
Is it worth switching to a cheaper model when one comes out?
You only know with your own set of cases to compare before and after. Changing the model name and deploying is easy, and the effect on the invoice shows up immediately; the effect on quality does not show up until customers complain. Without an evaluation set, the only thing you know for certain is that the bill changed.