An AI invoice arrives as a single figure, but it is the product of half a dozen separate decisions. Model choice, output length, hidden reasoning, caching and timing can each move the same workload's cost several times over.
An AI invoice is one number per vendor per month. It tells you what you spent. It tells you nothing about why, and so nothing about what would change it.
Underneath that number, every workload is the product of a handful of decisions, most of them made by a developer on the day the feature was built and never revisited. Each one can move the cost of the same work by a large multiple. Here they are in roughly the order they matter.
Which model does the work?
This is the biggest lever and the one least often pulled. The price gap between a vendor's flagship model and its smaller models is routinely an order of magnitude, and the gap across vendors is wider still. Many workloads, such as classification, extraction, routing and short summaries, do not need the flagship. They were put on it because it was the default, or because it was what the prototype used.
The test is not which model is cheapest. It is which is the cheapest model that does this particular job well enough, judged on your own examples. A workload moved from a flagship to a smaller model that passes the same test commonly costs a small fraction of what it did.
How long are the answers?
Output is priced at several times the rate of input across the major vendors. That makes the length of what comes back a cost decision, not a style choice. A system that returns a 600-word answer where 150 words would do is paying for four times the output it needs on every request.
Ask for structure instead of prose where the output feeds another system, set length limits that match the use, and stop asking for explanations nobody reads.
Is the model thinking out loud on your bill?
Reasoning models work through a problem before they answer, and that working is billed. OpenAI's documentation is explicit: reasoning tokens are not visible via the API but are billed as output tokens. Other vendors' extended thinking works the same way in principle.
That means two requests that return identical answers can cost very different amounts, and the difference does not appear in the response you can see. Reasoning earns its cost on genuinely hard problems. On a simple lookup it is overhead. Check which workloads have it switched on, and whether they need it.
Is the same context paid for again and again?
Many workloads send the same long instructions, policy text or reference document at the start of every request. Prompt caching lets the vendor reuse that repeated prefix instead of processing it afresh. At Anthropic, for example, reading from the cache is charged at a tenth of the normal input price, a 90% saving on the cached part, while writing to the cache costs a premium of 25% for a five-minute cache.
The arithmetic is worth doing per workload. If 80% of each request's input is a repeated prefix and it is read from cache at a 90% discount, the input cost of that workload falls by 72%. If the prefix changes every time, caching saves nothing and the write premium makes things slightly worse. The saving depends entirely on how much is genuinely repeated, and optimistic assumptions about cache hit rates are the most common reason a forecast turns out wrong.
Does it need an answer now?
Overnight reports, back-office document processing, evaluations and bulk classification rarely need a reply in seconds. The major vendors offer a batch route for exactly this work: requests are processed within a window of hours instead of immediately, at a substantial discount. Anthropic, for example, prices batch processing at 50% off both input and output.
Half price for the same output, for any workload that can wait, is the easiest saving on most bills. It is also the one most often missed, because the feature was built as an interactive call and nobody went back.
How often does the work actually run?
Volume is the term that multiplies all the others, and it is the one that grows without anyone deciding it should. Agents that call a model several times per task, retries on failure, and conversations that re-send their whole history at every turn all add requests that no business owner ever approved. Count requests per workload per month, and know which of them are retries.
Putting the levers together
The levers stack. Consider a nightly document summary job currently on a flagship model, with long free-text answers and no caching. Suppose testing shows a smaller model does the job to the same standard at a tenth of the price, the answers can be cut to a third of their length, a large shared instruction block can be cached, and the job can run as a batch.
Each step is modest on its own. Together they can take the cost of that job down by well over 90%, with no change to what the business receives. The point is not the exact figure, which depends on your workload. It is that the first number on the invoice was never a fixed cost of doing the work.
Key takeaways
- An AI invoice is the product of decisions, most of them made once and never revisited
- Model choice is the largest lever; test the cheapest model that does the job, on your own examples
- Output costs several times input, so answer length is a cost decision
- Reasoning tokens are billed even though you cannot see them
- Caching saves only on what genuinely repeats; batch processing halves the cost of work that can wait
- Volume multiplies everything; count retries and agent calls, not only user requests
The AI Spend Forecast prices each of your workloads against these levers, from your own bills, and shows the saving each one would bring before anything changes.