AI Cost

How to build an AI spend forecast your finance team will sign off

A step-by-step method for turning AI invoices into a forecast by workload: what to collect, how to measure each job, how to price the alternatives and how to present a budget that survives a finance review.

Most AI budgets are set by taking last month's invoice and adding a percentage. That works until it does not, and then nobody can explain why. A forecast that finance will sign off has to be built from the work itself: what each workload does, how often, and what it costs per run.

This guide sets out the method we use. You can follow it with a spreadsheet and your own exports.

Step 1: Collect what you already pay for

Gather two things for the last one to three months.

  • Usage or billing exports from every AI vendor you pay by the token. Every major vendor offers a usage export or a usage API; a CSV with honest column names is enough.
  • A list of every seat licence: vendor, plan, number of seats, price and which teams hold them.

Do not forget the indirect routes. AI bought through a cloud marketplace or bundled into another product's bill is still AI spend, and it is often the part nobody owns.

Step 2: Turn the invoice into workloads

An invoice is organised by vendor and model. A forecast has to be organised by workload: the business job that caused the spend, such as answering customer emails, summarising case files or classifying invoices.

Map each API key, project or workspace in the exports to the workload it serves. Where one key serves several workloads, split it, either by the request metadata you log or by agreeing a sensible allocation with the owner. From now on, give every new workload its own key or project. Attribution is far easier to build in than to reconstruct.

Step 3: Measure the shape of each workload

For each workload, record five numbers per month:

  1. Requests.
  2. Average input tokens per request, including any documents and conversation history sent with it.
  3. Average output tokens per request.
  4. The share of the input that is the same on every request, which is what caching could reuse.
  5. Tool calls per task, for agents and assistants that search or call other systems.

Use measured figures from the exports wherever they exist, and label anything you estimated. A forecast that mixes measured and assumed numbers without saying which is which will not survive its first question.

Step 4: Price each workload across the alternatives

For each workload, list the models that could do the job within your rules on data residency and approved vendors. Price the measured shape on each one, then apply the levers that fit:

  • Caching, on the share of input that genuinely repeats.
  • Batch processing, typically at half price, for anything that does not need an answer in seconds.
  • Shorter output, where the answers are longer than their use requires.

Then test before you recommend. Run a sample of real requests through the cheaper candidate and have the workload's owner judge the results. A cheaper model that fails the job is not a saving.

Write down the date of the price list you used. Prices change constantly, and a figure without a price date cannot be checked later.

Step 5: Decide seats against tokens

For each seat plan, price what the seat holders actually do at the same vendor's pay-per-use rates and find the break-even point. Expect a spread: heavy users for whom the seat is good value and light users for whom it is not. Record what each seat provides beyond the model, because a seat is also an interface, administration and data controls.

Step 6: Build the budget

Roll the workloads up into a monthly and annual run-rate, by workload and by vendor. Then show what it becomes at half and at double the expected volume, because volume is the assumption most likely to be wrong and finance will ask.

If there is a spending ceiling, show what it buys: which workloads fit inside it on which models, and what would have to give if volume grows.

Step 7: Present it so it can be signed

Open with the finding in one paragraph. Then one row per workload with the recommended model and the sentence that justifies it. Put caching and batch savings on separate lines so their assumptions can be challenged on their own. Label every figure as measured or estimated, and print the price date on the front.

Step 8: Re-forecast every quarter

Prices, models and volumes all move within a quarter. Re-run the forecast from the bill you just paid, compare it with what you predicted, and explain the difference. The first forecast is a starting point; the second one is where it becomes useful.

A checklist before you send it

  • Every workload has an owner and a business description
  • Every figure is labelled as measured or estimated
  • The price list date is on the front page
  • Caching and batch savings rest on stated assumptions
  • Cheaper models were tested on real requests before being recommended
  • The budget is shown at half and double volume

If you would rather not build it yourself, the AI Spend Forecast follows exactly this method over three weeks at a fixed price, from your own exports, and leaves you with the working forecast to re-run.

AI CostAI Spend ForecastHow-to guide

Want to apply this to a specific process?

Bring the workflow you had in mind. We will talk through whether these ideas apply to it, and what it would take to find out.

Around four minutes. Indicative guidance based on your answers.

Region & currency

Changes spelling, terminology, the data-protection regime named in our notices, and the currency used in indicative figures. It does not change where we are or quote you a price in your currency. ETT is headquartered in Dallas, with offices in London and Vancouver.