Part of our guide: The real cost of AI
Every budgeting conversation about AI starts in the wrong place, with somebody asking what the licence costs and getting an answer that is knowable. That answer goes into the plan and turns out to be the one number in the whole conversation that was never in question.
The problem with agentic AI is not that it costs more than a chat interface: it is that the same task, run twice, can cost thirty times more, and spending more does not buy a better answer.
The finding that breaks the budget line
An academic preprint published in April 2026, with authors including Erik Brynjolfsson and Alex Pentland, measured token consumption across agentic coding tasks, producing three results that matter to anyone signing off a business case.
Runs on the same task differed by up to 30 times in total tokens, and higher token usage did not translate into higher accuracy. Accuracy often peaked at intermediate cost and then saturated, and when models were asked to predict their own token cost, the correlation came out at 0.39, with systematic underestimation.
Agentic tasks consumed around a thousand times more tokens than code chat, with input tokens rather than output tokens driving the bill.
The first number is worth sitting with, because a thirty-times variance on an identical task is not a forecasting difficulty and is instead the absence of a forecastable unit. You cannot build a cost-per-transaction figure out of it, and so you cannot build a business case in the shape that finance departments expect.
Anthropic's own documentation gives smaller multiples for its systems, around four times more tokens for agents than chat and about fifteen times for multi-agent systems, without a stated sample or distribution. The two sets of figures are not interchangeable, and the difference between them is itself the point, showing that nobody has a stable number.
The market agrees, quietly
The evidence that this is systemic rather than a measurement quirk sits in the way organisations behave when they plan.
A survey of 396 organisations published in July 2026 found that only 11 per cent forecast their AI spend within 10 per cent. That holds despite 98 per cent tracking AI infrastructure costs and 95 per cent assigning formal AI budgets, and they still cannot predict the number they watch so closely.
The same research found the top source of unexpected cost was data platform usage overages at 47 per cent, ahead of model token costs at 43 per cent. The bill that people plan for is not the bill that surprises them.
Meanwhile the discipline has arrived at speed, with AI cost management now standing as the number one skills gap inside it. 98 per cent of FinOps practitioners now manage AI spend, up from 63 per cent in 2025 and 31 per cent in 2024. In June 2026 the FinOps Foundation and the Linux Foundation announced an intent to form a body for AI billing. That body would create open standards, with backing from Oracle, Google, Microsoft, Accenture, IBM, JPMorganChase, KPMG, Salesforce, SAP and ServiceNow.
A consortium of that size convening on billing standards is strong evidence that AI cost is not yet measurable in a comparable way.
Three levers that do not work the way people assume
Batch discounts are unavailable to agents. Anthropic's documentation states the 50 per cent Batch API discount does not apply to its managed agent sessions, because those sessions are stateful and interactive. There is no batch mode, and the cheapest lever on the price list is the one that agentic workloads cannot pull.
Provisioned throughput is often more expensive per token. The FinOps Foundation found one scale tier priced 67 per cent above standard rates and another priced 27 per cent above the standard. What you are buying there is latency and a service level, not savings.
Prices do not only fall. Google has published a dated 100 per cent increase on one model, from 0.75 to 1.50 dollars input and 3.75 to 7.50 dollars output, effective 1 January 2027. Long-context premiums exist on some providers and have been removed by others, and a three-year plan built on an assumption of continuous price decline carries an unhedged position.
The cost nobody has costed
For regulated firms there is a specific exposure in all of this, and it comes with no published benchmark at all.
Model providers deprecate models, and one provider commits to at least 60 days' notice before retirement while another commits to at least six months for generally available models. The first provider's 2026 record shows a model deprecated on 5 June and then retired on 5 August, a gap of sixty-one days. Nine model retirements happened at that provider between July 2025 and August 2026.
If your model risk governance rests on validation evidence for a specific model, a supplier can withdraw it with roughly two months of notice, and re-validation is not optional. There is no published cost or effort estimate for post-upgrade re-validation anywhere in the public record, and we have looked for one.
That absence should go in your risk register as it stands, because an unquantified obligation with a two-month trigger is a materially different thing from an annual licence.
The Bank of England and PRA heard the same from firms in roundtables with banks and insurers, where second-line risk functions are approaching AI with caution in ways that may delay deployment. Firms said their traditional model risk management approach to validation would not be sustainable in its current form as generative and agentic systems proliferate.
Where the savings actually go
The Bank of England's Agents, reporting in July 2026 from conversations through to the end of June, found that AI cost savings "are being offset in part by rising expenditure on software, cloud services and AI licences."
This is the mechanism behind many disappointing business cases, where the saving is real and lands in one budget and the cost is real and lands in another. If those two lines sit with different owners, the programme reports success and the P&L stays exactly where it was.
The fix is unglamorous and it works: putting one owner and one line in place, reporting gross rather than net and reviewing the position monthly.
How to budget anyway
You cannot forecast the unit cost, and you can still control the total.
Set a ceiling before a forecast. Where the unit is unstable, budget the work as a capped consumption line with hard alerting rather than as a projection. This is how organisations already handle cloud egress, and that remains the closest existing analogue for an unstable unit cost.
Measure your own variance early. Run the same representative task twenty times inside your own environment and look at the spread of results before you commit. Your distribution is the only one that matters for your own business case, and it takes a day to produce.
Instrument before you scale. A single user interaction in an agentic system may produce five, ten or fifty separate model calls underneath the interface. If you cannot attribute cost to a process, you cannot manage it, and retrofitting attribution after go-live is materially harder.
Budget the run, not the build. The lines that surprise organisations are data platform overages, re-validation after forced model changes and the licence and cloud spend that grows alongside the saving, and none of them appear in a pilot.
Do not put a fake precision in the paper. If the real position is that the unit cost varies by an order of magnitude, say so and manage it with a cap. A board can work with a bounded uncertainty and it cannot work with a confident number that turns out to be fiction.
We build this picture as part of a Breathe discovery sprint, and we wrote about the wider cost structure in the real cost of AI for a COO. If your AI business case currently carries a single figure for the run cost, it is worth a second look.