Skip to content
Decision support

AI is no longer a licence cost. It is a consumption nobody owns

A licence per person is predictable by construction. A consumption cost is unpredictable for the same reason. Most organisations budgeted the second in the mould of the first, and find out once the money is already spent. The question is not which model is cheapest, but what a finished task costs in your organisation and who owns that line.

Andreas Olsson9 min read

Two identical stone plinths. On the left one a block sits squarely on top, on the right an identical block has spilled over the edge and across the floor.

Key insights


  • Uber spent its entire 2026 AI budget by April. The cause was not overuse but a consumption cost budgeted as a licence cost.
  • Across 700 organisations with at least 1,000 employees, 52 percent say the AI cost has no clear owner and 72 percent have had a surprise bill in the past year.
  • Output from GPT-5.5 costs 30 dollars per million tokens. The same output from DeepSeek V4 Flash costs 0.18. That is roughly a hundred and sixty-seven times.
  • DeepSeek's share of token flow through OpenRouter went from just under ten percent to eighteen in five months, while the capability gap to the frontier has held at three to six months for eighteen months.
  • Cheapest per token is not cheapest per outcome. A weaker model that needs three attempts at the same task costs more, not less.

A budget that ran out in April

In May, Forbes reported, citing The Information, that Uber had spent its entire 2026 AI budget by April. The company's CTO, Praveen Neppalli Naga, confirmed it, and described the position as being back at the drawing board on the assumptions.

The figures behind the headline are more instructive than the headline. Around 5,000 engineers, and usage that reached 95 percent monthly through the spring. The average ran at 150 to 250 dollars per engineer per month. Heavy users ran between 500 and 2,000. The CTO himself spent 1,200 dollars in a single two-hour session.

It is easy to read that as waste. It is harder, and more useful, to read it as a budgeting error of a kind that will repeat in every organisation whose adoption succeeds.

Why it was not waste

Until recently software had a price per person per year. Finance counted the people, multiplied by an amount and got a number that held. It could be wrong if more people were hired, but it could not be wrong by a factor of three.

A consumption cost does not behave that way. It is set by how much work is actually done, and it grows with the success of the adoption. The better it goes, the more it costs. That is the opposite of a licence, where the cost is incurred whether the tool is used or sits idle.

The only AI budget that holds all year is the one nobody in the organisation used.

Uber's forecast was not careless. It was made in the right template for the wrong shape of cost. And ninety-five percent adoption, the number every leadership team says it wants, was what triggered the overrun.

The line nobody owns

Ask who in your organisation owns the AI cost, and the answer is likely to be a pause followed by three half-answers.

IT owns the licences, because that is where the contracts sit. The business owns the usage, because that is where the work happens. Finance owns the budget, because that is where the number lands. The consumption arises in the gap between the three, and a gap has no manager.

Uber is a single case. How common it is for the line to have no owner has been measured. In Harness's 2026 State of AI in FinOps, 52 percent say there is no clear owner of the AI cost. 72 percent have had an unexpected cost spike or a surprise bill in the past twelve months, and 56 percent say AI budgeting is guesswork.

The base is 700 engineering leaders and practitioners at organisations with at least 1,000 employees that actively use AI or LLM services and carry a recurring AI cost. 300 respondents in the United States and 100 each in the United Kingdom, France, Germany and India. Fieldwork ran from May to June 2026, collection was an online survey, and no margin of error is stated.

What 700 organisations say about their AI cost

52%
No clear owner of the AI cost
Accountability split across engineering, FinOps, finance and IT.
72%
Unexpected cost spike or surprise bill
In the past twelve months.
56%
Say AI budgeting is guesswork
700 respondents at organisations with at least 1,000 employees.

Source: Harness, 2026 State of AI in FinOps, May to June 2026

Three things belong with those figures. The survey is published by a vendor that sells tooling for cost control, which is reason to read it as an order of magnitude rather than an exact measurement. The sample is large organisations in five countries and holds no Nordic base. And the survey does not measure the difference between a licence and a consumption. That question is not put in the questionnaire, and the figures therefore do not carry it.

It is the same pattern that governs the competence question: a system nobody owns is used anyway, only without anyone making the decisions about it. The difference is that the cost shows up in a report sooner or later, and it shows up for the finance director before it shows up for anyone else.

What the consumption has returned so far

31%
Have AI in daily use across the organisation
Nordic organisations, up from seven percent the year before. Tietoevry 2026.
22%
Say the return has met or exceeded expectations
ISACA AI Pulse Poll 2026, more than 3,400 respondents.
4%
Say AI is a critical part of core infrastructure
Nordic organisations. Tietoevry 2026.

Source: Tietoevry Nordic AI Survey 2026 and ISACA AI Pulse Poll 2026

Those figures are uncomfortable in this particular context. Usage has grown sharply in a year, while the share saying the return has met expectations is barely a fifth and the share calling AI critical to core infrastructure is four percent. The cost has scaled faster than the value can be shown, and that is exactly the situation in which a clear owner of the cost is worth most.

The price per token is not the price per task

At the same time the pricing picture has pulled apart in a way that makes the choice of model a finance question rather than a technology one.

OpenRouter, which routes traffic between models and therefore sees both prices and volumes at close range, set two price points side by side on 30 June 2026. DeepSeek V4 Flash on its cheapest endpoint costs 0.09 dollars per million tokens in and 0.18 out. GPT-5.5 costs 5 in and 30 out.

That is not a margin. It is roughly a hundred and sixty-seven times on output, for work that in many cases is interchangeable.

The shift also shows in the volumes and not only in the price list. DeepSeek accounted for just under ten percent of token flow through OpenRouter at the start of the year. By the start of June the share was eighteen percent, a doubling in five months. Among consumer users, close to a third of all tokens already go to DeepSeek models, and through 2025 American models accounted for around three quarters of everything used. The direction is measured, not guessed.

Price per million output tokens, June 2026

0.2dollars
DeepSeek V4 Flash
Cheapest endpoint. Input 0.09 dollars.
1.2dollars
MiniMax M3
Weighted average across providers.
3.3dollars
GLM 5.2
Weighted average across providers.
30dollars
GPT-5.5
Input 5 dollars.

Source: OpenRouter, June 2026

Two orders of magnitude between the extremes means the same task can cost a hundred times more depending on who performs it. For a single question in a chat window that does not matter. For an agent running thousands of times a month it is the difference between a line in the budget and an item in the board report.

And the premium is harder to justify than it was. OpenRouter describes the distance between open-weight models and the leading labs as a steady three to six months, and notes it has held there for eighteen months. That is not a gap closing, but it is not a gap widening either.

Not which model is cheapest, but what a finished task costs in your organisation.

There is an opposing effect that weighs more heavily. A weaker model that needs three attempts at the same task costs more than a costlier one that manages it in a single pass, and it also costs time for whoever has to check the result. That comparison cannot be read off a price list. It has to be measured in your own operation, on your own tasks.

Three things to do before the next budget

None of them waits on anything.

Measure cost per outcome on your three most common tasks. Not per user and not per month, but what it costs to get one thing finished. That is the number that can be compared between models, and it is the number missing from almost every discussion of AI cost.

Give the line a named owner. A person in the business who can see the cost, decide which model may be used for what, and say no. A function is not enough, because it is between functions that the cost arises.

Decide where the ceiling sits and what happens when it is reached. A ceiling with no agreed consequence is a number in a document. Decide in advance whether the work stops, moves to a cheaper model, or needs someone to approve an increase, and who that someone is.

The part that is not a cost question

The choice of model has one more dimension, and it should not be confused with the price.

Where the processing happens, under whose jurisdiction it happens, and what becomes of what you put in are questions with their own answers and their own consequences. They are not settled by what a million tokens costs. For a European organisation handling personal data, sensitive material or procurement documents, that is the question that sets the frame, and the cost comparison is then made inside that frame.

The order matters. Choose a model on price and discover the frame afterwards, and the work is done twice. Set the frame first and there are fewer options and an easier decision.

The full EU AI Act calendar, every date and what applies from when, sits on our EU AI Act page.


Common questions

Because the shape of the cost changed and the budgeting template did not. A licence per user is predictable: headcount times a price. A consumption cost is set by how much work is actually done, which nobody can know in advance and which grows precisely when adoption succeeds. Uber spent its entire 2026 AI budget by April, and the cause was not waste but a forecast built on the wrong kind of assumption.

There is no answer per user, only per task. At Uber the average was 150 to 250 dollars per engineer per month, while heavy users ran between 500 and 2,000 and a single two-hour session cost 1,200 dollars. The spread inside one organisation is wider than the difference between most software licences, and that is the point: an average per user does not describe the cost.

The price per token is what the vendor charges. The cost per task is what it actually takes to get something finished: price per token times tokens used times the number of attempts required. A cheaper model that has to be run three times on the same task is more expensive than a costlier one that completes it on the first attempt. Cost per task is the only comparison that matters, and it can only be measured inside your own operation.

Per token, yes, and the gap is large. According to OpenRouter's June 2026 data, DeepSeek V4 Flash on its cheapest endpoint costs 0.09 dollars per million tokens in and 0.18 out, while GPT-5.5 costs 5 in and 30 out. That is roughly a hundred and sixty-seven times on output. The same source shows DeepSeek's share of token flow rising from just under ten percent to eighteen over the first five months of the year. But price per token is not price per outcome, and for a European organisation the choice of model is also a question of where processing happens and under whose jurisdiction, which is not a cost question at all.

The gap is measured and it is smaller than the price difference suggests. OpenRouter describes the distance between open-weight models and the leading labs as a steady three to six months, and notes it has held there for eighteen months. For a task where half a year of lag in capability makes no difference, the premium is hard to justify. For a task where it does, the premium is easy to justify. That is decided per task, not per organisation.

A named person in the business, not a function. Today IT owns the licences, the business owns the usage and finance owns the budget, while the consumption arises in the gap between them. Whoever owns the line needs to see the cost per task, decide which model may be used for what, and settle what happens when a ceiling is reached.

Three things. Measure what your three most common AI tasks cost per finished outcome rather than per user. Give the cost line a named owner. And decide in advance where the ceiling sits and what happens when it is reached, because a ceiling with no agreed consequence is just a number in a document.


If this lands on your desk, we should talk.

Ampliro Insights

New analysis, roughly weekly.

We write when the rules change and when something turns out to work in practice. One piece at a time, no sequences, and you can leave from any issue.

We store your address to send Ampliro Insights, and for nothing else. More in the privacy policy.