For about thirty years, business software was bought like a gym membership. A fixed amount per employee per month. You knew the bill a year in advance and it made no difference whether someone used the software twice a day or never opened it.

AI is not sold that way. It is sold like a taxi meter. Every question asked, every document read, every task an agent picks up on its own consumes small units of computing work called tokens, and every token carries a price. So the invoice is no longer set by how many people you employ. It is set by how much work those people hand over, and by how much of that work the software chooses to do before it answers.

That reads like a change in a purchasing model. It is closer to a change in what kind of business you are running, because a fixed cost and a variable cost are managed by different people, on a different cadence, with different questions in the room. Most organisations have made the switch without noticing that the switch happened.

Ask four executives in one room what their AI will cost next year and you will get four numbers. All four will be licence numbers. None of them will be wrong, exactly. They will all be answers to a question that stopped being the relevant one somewhere around the middle of this year.

What the seat was quietly doing for you

The old model was never really about software. The seat was a proxy and a good one, standing in for how much value a company was pulling out of a system. Stable enough to forecast a year out. Legible enough for a CFO to sign without reading twice.

Everything downstream got calibrated on that proxy. Enterprise agreements, discount ladders, approval thresholds, the way software companies are valued on seat expansion, sales compensation, partner margin, how a CIO’s performance is judged.

And it was doing one more thing that nobody thanked it for. It was capping exposure.

An employee has eight hours in a day. That ceiling was never a policy, it was a fact, and it kept software spend bounded by something no vendor could influence. Licences could be added. A thousand of them could not appear overnight by accident.

Then the unit changed. Agentic AI is priced by consumption and the seat has been demoted to an access fee. Microsoft 365 Copilot still lists at thirty dollars per user per month, but Cowork, Copilot Studio and the Work IQ API all run on Copilot Credits at a published cent apiece and Microsoft’s own documentation states that licences “act as an entry point enabling access to AI services billed on a pay-as-you-go basis.” Anthropic is blunter: Claude Enterprise is twenty dollars a seat, and “usage is billed as you go at API rates, based on what your team uses.” OpenAI is replacing its older model-and-message-count method with dollar-per-token billing for new enterprise agreements. Google bundled Gemini into Workspace seats, then launched a Gemini Enterprise edition with a zero dollar seat price and usage-based billing only.

IDC expects pure seat-based pricing to be obsolete by 2028, forcing seventy percent of vendors to refactor their value proposition.

Agents, for their part, do not get tired. They have a schedule and a retry loop. Anthropic’s documentation notes that agent teams consume around seven times the tokens of a standard session. Jellyfish measured per developer token consumption rising roughly nineteenfold in nine months.

So the natural ceiling came out of the model and almost nobody replaced it with a deliberate one.

The transfer nobody negotiated

Consumption pricing is not mainly a pricing decision. It is a transfer of variance and it happened without most buyers registering it as a change in their risk position.

Sanchit Vir Gogia of Greyhound Research puts it about as directly as an analyst can: vendors are “transferring the cost volatility of AI compute to customers while monetizing customer-side productivity gains as margin.” Under seat pricing the vendor carried the compute risk. Expensive inference was their margin problem, cheap inference was their upside. That risk now sits with the buyer, on an input the buyer neither owns nor controls.

And it genuinely is not under the buyer’s control, which is the part that gets missed in negotiation. In a consumption contract the vendor sets the speed of the meter. A change to the default model, to routing, to how much context gets resent, to how many times an agent retries, all move the invoice without one line of the contract changing.

Microsoft demonstrated that lever on itself. Alongside giving its divisions token budget targets from July 2026, it moved staff onto a cheaper default model in GitHub Copilot. Deliberate cost control, pulled in the direction the payer wanted. Nothing guarantees the same lever gets pulled the same way when the payer is somebody else.

The asymmetry that kills programmes that were working

KPMG’s Global AI Pulse for the second quarter of 2026 surveyed 2,145 senior leaders. Twenty-four percent had scaled back or narrowed a deployment of AI agents. Another twenty-five percent had delayed or paused further rollout. The stated trigger was the same in both cases: expected cost started to outweigh the value generated. The easy conclusion is that the technology underdelivered. The evidence points somewhere else.

Look at how the two sides of the ledger arrive. Cost arrives centralised, itemised and monthly. One invoice, one owner, painfully legible. Benefit arrives decentralised and unlabelled. Ninety minutes here. A handover that no longer needed a meeting. A report nobody stayed late for. Spread across forty teams and attributed to none of them.

Centralised visible cost against decentralised invisible benefit is a fight the AI programme loses every time, whether or not it is creating value. Any CFO alive will cut the line item they can see before the value they cannot.

Which inverts the usual advice. The instinct is to instrument cost first, because cost is the number that hurts. The better move is to instrument benefit first, because benefit is the number that is missing. Build only the cost ledger and a programme that was working will eventually be cancelled while looking like one that was not.

The programmes stalling in 2026 are largely not failing on disappointing technology. They are failing on unbudgeted success.

For about thirty years, business software was bought like a gym membership. A fixed amount per employee per month. You knew the bill a year in advance, and it made no difference whether someone used the software twice a day or never opened it.

AI is not sold that way. It is sold like a taxi meter. Every question asked, every document read, every task an agent picks up on its own consumes small units of computing work called tokens, and every token carries a price. So the invoice is no longer set by how many people you employ. It is set by how much work those people hand over, and by how much of that work the software chooses to do before it answers. That reads like a change in a purchasing model. It is closer to a change in what kind of business you are running, because a fixed cost and a variable cost are managed by different people, on a different cadence, with different questions in the room. Most organisations have made the switch without noticing that the switch happened.

Ask four executives in one room what their AI will cost next year and you will get four numbers. All four will be licence numbers. None of them will be wrong, exactly. They will all be answers to a question that stopped being the relevant one somewhere around the middle of this year.

What the seat was quietly doing for you

The old model was never really about software. The seat was a proxy, and a good one, standing in for how much value a company was pulling out of a system. Stable enough to forecast a year out. Legible enough for a CFO to sign without reading twice.

Everything downstream got calibrated on that proxy. Enterprise agreements, discount ladders, approval thresholds, the way software companies are valued on seat expansion, sales compensation, partner margin, how a CIO’s performance is judged.

And it was doing one more thing that nobody thanked it for. It was capping exposure.

An employee has eight hours in a day. That ceiling was never a policy, it was a fact, and it kept software spend bounded by something no vendor could influence. Licences could be added. A thousand of them could not appear overnight by accident.

Then the unit changed. Agentic AI is priced by consumption and the seat has been demoted to an access fee. Microsoft 365 Copilot still lists at thirty dollars per user per month, but Cowork, Copilot Studio and the Work IQ API all run on Copilot Credits at a published cent apiece, and Microsoft’s own documentation states that licences “act as an entry point enabling access to AI services billed on a pay-as-you-go basis.” Anthropic is blunter: Claude Enterprise is twenty dollars a seat, and “usage is billed as you go at API rates, based on what your team uses.” OpenAI is replacing its older model-and-message-count method with dollar-per-token billing for new enterprise agreements. Google bundled Gemini into Workspace seats, then launched a Gemini Enterprise edition with a zero dollar seat price and usage-based billing only.

IDC expects pure seat-based pricing to be obsolete by 2028, forcing seventy percent of vendors to refactor their value proposition.

Agents, for their part, do not get tired. They have a schedule and a retry loop. Anthropic’s documentation notes that agent teams consume around seven times the tokens of a standard session. Jellyfish measured per-developer token consumption rising roughly nineteenfold in nine months.

So the natural ceiling came out of the model, and almost nobody replaced it with a deliberate one.

The transfer nobody negotiated

Consumption pricing is not mainly a pricing decision. It is a transfer of variance, and it happened without most buyers registering it as a change in their risk position.

Sanchit Vir Gogia of Greyhound Research puts it about as directly as an analyst can: vendors are “transferring the cost volatility of AI compute to customers while monetizing customer-side productivity gains as margin.”

Under seat pricing the vendor carried the compute risk. Expensive inference was their margin problem, cheap inference was their upside. That risk now sits with the buyer, on an input the buyer neither owns nor controls.

And it genuinely is not under the buyer’s control, which is the part that gets missed in negotiation. In a consumption contract the vendor sets the speed of the meter. A change to the default model, to routing, to how much context gets resent, to how many times an agent retries, all move the invoice without one line of the contract changing.

Microsoft demonstrated that lever on itself. Alongside giving its divisions token budget targets from July 2026, it changed the default model its own staff use in GitHub Copilot, to one priced below the model it replaced on Microsoft’s published enterprise rate card. Deliberate cost control, pulled in the direction the payer wanted. Nothing guarantees the same lever gets pulled the same way when the payer is somebody else.

The asymmetry that kills programmes that were working

KPMG’s Global AI Pulse for the second quarter of 2026 surveyed 2,145 senior leaders. Twenty-four percent had scaled back or narrowed a deployment of AI agents. Another twenty-five percent had delayed or paused further rollout. The stated trigger was the same in both cases: expected cost started to outweigh the value generated.

The easy conclusion is that the technology underdelivered. The evidence points somewhere else.

Look at how the two sides of the ledger arrive. Cost arrives centralised, itemised and monthly. One invoice, one owner, painfully legible. Benefit arrives decentralised and unlabelled. Ninety minutes here. A handover that no longer needed a meeting. A report nobody stayed late for. Spread across forty teams and attributed to none of them.

Centralised visible cost against decentralised invisible benefit is a fight the AI programme loses every time, whether or not it is creating value. Any CFO alive will cut the line item they can see before the value they cannot.

Which inverts the usual advice. The instinct is to instrument cost first, because cost is the number that hurts. The better move is to instrument benefit first, because benefit is the number that is missing. Build only the cost ledger and a programme that was working will eventually be cancelled while looking like one that was not.

The programmes stalling in 2026 are largely not failing on disappointing technology. They are failing on unbudgeted success.

Where this goes next

Everything above is observable today. What follows is a forecast, with the mechanism attached to each one, stated specifically enough to be held against the record later.

The falling price of tokens will raise the bill, not lower it

This is the objection in every room. Model prices drop constantly, so cost governance is a temporary problem that the technology curve will solve.

It will not, and the evidence is already in. Per-token prices have fallen hard, and in the same window per-developer consumption rose roughly nineteenfold in nine months. Anthropic, OpenAI and Google all price cached context at about a tenth of fresh input and halve the bill again for batch work, so the efficient path is already dramatically cheaper than the default path, and total spend still climbs. Goldman Sachs projects global token usage growing twenty-four times by 2030.

Cheaper units do not reduce consumption of something whose usefulness scales with how much of it gets used. They expand it. Jevons wrote that down in 1865 about coal, and it has held for bandwidth, for storage and for compute since.

The prediction, then, is the uncomfortable one. Organisations currently waiting for prices to fall far enough to make this a non-issue will spend more in 2028 than in 2026, on cheaper tokens and will have lost two years of learning to competitors who started measuring in 2026.

The energy buyer already knows how to buy AI, and nobody has told them

Look at the shape of AI pricing as it now stands, rather than at the individual numbers.

There is a spot rate, which is pay-as-you-go per token. There is a forward market: Microsoft’s Copilot Credit pre-purchase plans commit a buyer for a year at five to twenty percent off, and unused units expire, which is a take-or-pay contract wearing a software name. There is off-peak pricing, in the form of batch processing at half price for work somebody is willing to defer. There is peak pricing, in the form of Google’s priority tier at 1.8 times standard rate for lower latency. Google charges cache storage by the hour, which is a holding cost.

Spot, forward, take-or-pay, peak, off-peak, storage. That is not a software price list. That is an energy market, and it arrived fully formed while enterprise procurement was still asking for a bigger discount.

Which locates the scarce skill precisely. The competence needed to buy a volatile metered input is hedging, not haggling, and it has existed in utilities and commodities for fifty years: forward commitments, collars, tariff structures, load shifting, price caps.

Prediction: within two years, large enterprises will begin hiring AI procurement out of energy and commodity trading rather than out of software procurement and the first to do it will show materially lower cost volatility than peers holding better headline discounts. The cheapest possible head start is to walk down the corridor to whoever negotiates the electricity contract and show them a model rate card. They will recognise the structure on sight.

Vendors will discover that certainty carries a better margin than compute

Buyers dislike variance more than they dislike cost. That is not a new insight, it is the entire history of insurance.

Which points to the obvious product gap: a price cap. Not a discount, a guarantee. A committed band, an overage ceiling, an undertaking that a change to the default model will not reprice the year. A vendor with a large enough portfolio can pool that variance across thousands of customers far more cheaply than any single customer can absorb it alone, and can charge a premium for absorbing it.

Prediction: within two to three years the sharpest competitive weapon in enterprise AI will not be capability or per-token price, it will be a credible spend guarantee and the margin on selling certainty will beat the margin on selling compute. Expect it to appear first as a contract term rather than as a product and expect actual insurers to take an interest in this market.

There is an implication for buyers worth sitting with. If certainty becomes something vendors charge for, then variance is a cost currently being absorbed on the vendor’s behalf, for free. The customer is underwriting the supplier.

Token budget becomes a management right, then a hiring negotiation

Microsoft gives its divisions token budget targets and lets individual engineers see their own spend, reportedly running from hundreds to several thousand dollars a month. That is the first move in a sequence that is easy to extrapolate and rarely discussed.

Faros AI found the heaviest token users were roughly twice as productive as their peers while consuming ten times the tokens. Once that relationship becomes visible inside a company, a token allowance stops being a cost line and becomes a productivity input that people compete for, like headcount or equipment. Managers will hoard it, because it makes their teams look good. The real power map of the organisation shifts toward whoever controls the quota, which in most places is a person nobody has appointed.

Prediction: within eighteen to twenty-four months, AI budget appears in senior hiring conversations the way headcount, tooling and travel budget do today. Somebody will decline an offer because the allowance was too small, and that will be the moment this stops being read as an IT cost topic.

Measurement bias, not capability, will decide the order in which work gets automated

This is the prediction that runs hardest against the public conversation about AI and jobs.

If AI cost is legible per task and human cost is not, then AI looks expensive exactly where work is well measured, and free where it was never measured at all. A contact centre that knows its cost per resolved contact to the cent can hold an agent against a person and will sometimes find the agent unfavourable. A strategy function that has never costed a single deliverable will never generate a number that makes AI look bad.

So the automation frontier will not follow difficulty. It will follow accounting. Well-instrumented operational work gets scrutinised hardest and, paradoxically, defended best, because it can produce a comparison. Unmeasured senior knowledge work goes last, not because it resists automation but because nobody can price the alternative.

Prediction: the popular thesis that AI comes for knowledge work before operational work will invert in practice across the next five years and the reason will turn out to be management accounting rather than model capability. Any workforce strategy built on capability alone is modelling the wrong variable.

Consumption pricing is a waypoint, not the destination, and both sides are flying blind

Consumption pricing solves the vendor’s problem and creates the buyer’s, so it will not hold. Pressure will push toward outcome pricing, because outcome pricing sends the variance back across the table: pay per resolved ticket, per closed claim, per approved invoice.

The reason that has not happened yet is neither technical nor contractual. It is that vendors cannot price it either. Outcome pricing requires knowing your own cost per outcome, and almost nobody on either side of the table has that number today.

Prediction: the first credible vendor to publish genuine outcome pricing on a mainstream enterprise workload will take share far faster than its capability advantage justifies and will force the market inside two renewal cycles. On the buyer’s side, organisations that know their own cost per outcome first will negotiate those contracts from a position competitors cannot match, because they will be the only party in the room able to judge whether the price is good.

Cloud arguably ran this play already. The reading here is that between roughly 2012 and 2015 the companies able to price per transaction did not merely run leaner, they priced into competitors’ markets and were not followed, because the competitor could not work out whether the deal was profitable. The same asymmetry is loading now, and it rewards whoever measures first rather than whoever adopts hardest.

A role that does not exist today becomes a route to the top

Seat pricing made the CIO the buyer. Consumption pricing makes the CFO the payer and the engineer the spender, and leaves the middle empty.

That empty middle is where the money now moves. Cached context reads at a tenth of input price. Batch halves the bill. Model choice shifts unit cost five times over or more between a vendor’s lightest and heaviest tiers. Anthropic states plainly that unexpectedly high spend “usually traces back to long sessions that were never cleared or to Opus left as the default model.” No negotiated discount competes with those numbers, which means the largest single lever on an AI cost base is an architecture decision taken by an engineer on a Tuesday afternoon with no finance conversation anywhere near it.

The FinOps Foundation, which spent years building this discipline for cloud, now names cost per token and cost per inference as the unit metrics to build on, and reports that ninety-eight percent of its respondents manage AI spend today against thirty-one percent two years ago. Its sharpest line is about ownership rather than measurement: “A token quota is simultaneously a financial control and an engineering control, and it only works if both sides set it together.”

Prediction: within twelve to twenty-four months a named role appears between the CFO and the CTO, owning cost per outcome across the portfolio and it becomes one of the faster routes into an operating or finance leadership seat, much as cloud FinOps quietly produced a generation of infrastructure leaders. For anyone deciding where to place their strongest generalist this year, that is the seat.

What would change the argument

Predictions without failure conditions are opinions wearing confidence, so here are the conditions. If a major vendor returns to genuinely all-inclusive seat pricing at scale and holds margin doing it, the risk transfer argument weakens considerably. Google bundling Gemini into Workspace seats is a real counterexample, and the open question is whether it survives or whether the zero dollar Gemini Enterprise seat with metered usage turns out to be the direction of travel. If inference costs collapse by two orders of magnitude rather than one, part of this becomes a rounding error rather than a discipline. And if agent consumption plateaus once novelty passes, instead of compounding, the ceiling problem partly solves itself.

None of those look likely. They are simply the first places to check if this reading turns out to be wrong, and naming them is cheaper than pretending the argument has no soft edges.

The part that has not changed

Strip the tooling away and the shift is easy to state and hard to absorb. Enterprise technology moved from a fixed cost to a variable one, and variable cost businesses are run differently from fixed cost ones.

There is a version of this argument in which finance is the villain. The CFO as brake, the budget as the thing that kills the pilot, governance as the enemy of momentum. That reading is backwards. An unbudgeted AI programme is not a fast programme, it is a fragile one, and the KPMG numbers show precisely what happens to fragile programmes on the day the invoice lands. Finance in the room is what allows ambition to survive contact with the P&L.

Microsoft, of all organisations, made the point on itself. When its internal token guidance surfaced in August, executive vice president Jay Parikh’s framing was quoted everywhere: “Tokenmaxxing is not what we are optimizing for.” That reads as honesty rather than embarrassment. The company with the deepest commercial interest in unlimited consumption looked at its own bill and installed a meter.

The seat used to answer the budget question on everyone’s behalf. It does not any more, and nothing has replaced it except a decision that somebody has to own.

Which is the whole thing, really. Adding AI is easy. Redesigning work is leadership.

A note on the numbers

Vendor pricing in this area changes monthly. Every price and product mechanic above was checked against the vendor’s own published pricing or documentation in early September 2026. The research and survey figures, including the consumption growth measured by Jellyfish, the productivity comparison from Faros AI, the Goldman Sachs projection and the IDC forecast, come from named reporting rather than from primary publications, and are attributed in the sources below. Anything derived by arithmetic rather than read off a page is described that way in the text. One widely repeated claim, about a company spending half a billion dollars in a month on unmanaged licences, is deliberately absent: it traces to a single sentence from one unnamed consultant about one unnamed client, and it does not stand up.