Foundry4

AI and automation 8 min read

The second year is where AI automation gets expensive

Enterprise AI has moved from per-seat licences to per-action metering. That changes who controls your bill, and the vendor can redefine the unit you budgeted in.

On 1 September 2025, Microsoft changed the unit in which Copilot Studio is billed. The product’s documentation records the change plainly: the common currency for agents changed from messages to Copilot Credits, with no change to the quantity in a prepaid pack or to the pay-as-you-go rate.

Nobody’s bill moved on the day. What moved was the definition. Any forecast built on messages had to be rebuilt on credits, and the conversion between them is set by a rate card the vendor maintains and can revise. That is the single most under-appreciated fact about the current generation of enterprise AI: the buyer is no longer forecasting a quantity they control, but a derivative of a definition they do not.

The arithmetic, from the published numbers

Microsoft’s UK page lists Copilot Studio as tenant-wide packs of 25,000 credits at £153.80 per pack per month, exclusive of VAT, with a pay-as-you-go meter available through an Azure subscription and a discount of up to 20% for committing in advance. Microsoft 365 Copilot is listed separately at £23.10 per user per month on an annual commitment.

Divide the pack and each credit costs a little over 0.6p. Apply that to the published billing rates and the per-event costs fall out:

  • a hand-authored classic answer, 1 credit, about 0.6p
  • a generative answer, 2 credits, about 1.2p
  • an agent action, which includes triggers, deep reasoning and topic transitions, 5 credits, about 3.1p
  • tenant graph grounding for a message, 10 credits, about 6.2p
  • a hundred agent flow actions, 13 credits, about 8p for the hundred
  • a page through the content processing tools, 8 credits, about 4.9p
  • a premium generative voice minute, 75 credits, about 46p

Those figures are arithmetic on published list prices, not a quotation for any particular organisation, and enterprise agreements move them. But the shape is the point. Consider a body of a thousand staff each putting five grounded questions a day to an internal agent across twenty working days. That is a hundred thousand generative answers, two hundred thousand credits, eight packs, and a little over £1,200 a month. Turn on tenant graph grounding for the same volume and the grounding alone adds ten credits an interaction, five times the answer itself.

Nobody signs off a project on £1,200 a month. That is exactly the problem. The number is small enough to approve without a business case and elastic enough to multiply by ten when usage goes the way the sponsor hoped it would.

The seat licence and the meter interact, and that is where the money is

The rate card carries a second column that most summaries omit. Where the person using an agent is licensed for Microsoft 365 Copilot and the agent runs under that user’s identity, the same activities are shown as no charge. Classic answers, generative answers, agent actions, tenant graph grounding and the AI tools are all zero-rated in that scenario, subject to fair usage limits and to specific conditions on how agent flows are triggered. Computer-using agents are called out as not included.

That single design decision explains most enterprise buying behaviour this year. At £23.10 per user per month, a thousand-seat deployment of Microsoft 365 Copilot is around £277,000 a year before VAT, and it converts a large and volatile metered cost for internal use into a fixed and forecastable one. Buyers are not choosing seats because seats are cheaper per interaction. They are choosing them because a fixed number can be approved once and a variable one has to be defended every quarter.

The consequence is that the metering bites hardest exactly where the internal seat licence does not reach: agents facing customers on a website or in a contact centre, agents serving staff who are not licensed, and anything that drives a user interface. Those are also the deployments with the highest volumes and the least predictable demand, which is an unfortunate combination. An organisation that models its costs on internal pilots and then launches externally will discover the difference in the first full month of live traffic, not before.

Success is the cost event

Under a per-seat licence, the worst commercial outcome is that nobody uses the thing you bought. Under per-action metering it is inverted. Adoption is the cost curve. A campaign to get colleagues using the assistant is a campaign to increase the invoice, and the person running it usually has no visibility of the meter.

Two published terms turn that from a budgeting inconvenience into an operational risk.

The first is that capacity does not roll forward. Copilot Studio enforces purchased capacity monthly and unused Copilot Credit don’t carry over to the next month. An organisation with seasonal demand, which is most of them, therefore either overbuys for eleven months or runs short in the twelfth.

The second is what happens when you run short. Microsoft’s documentation states that if usage exceeds purchased capacity, technical enforcement applies and can result in service denial, and that for agent flows specifically, exhausted prepaid capacity blocks new runs. So the failure mode of a metered automation is not a larger bill. It is that the automation stops, on a day when demand was unusually high, which is the day it was least likely to be missed quietly.

Any organisation putting a metered agent into a process that has a service level attached to it needs to have decided in advance which of those two it prefers, and to have written the answer into the operating procedure. Very few have.

The three lines that never appear in the pilot

Licences are the forecastable part. The lines that arrive later are the ones a pilot cannot generate, because a pilot does not last long enough.

Evaluation. A rules-based automation is verified by a test suite that runs in seconds and either passes or does not. A system with a model in it has an error rate rather than a state, and establishing that rate requires a held-out set, a sampling regime and somebody senior enough to adjudicate the disagreements. That work recurs every time the underlying model is updated, and the update schedule belongs to the vendor. This is a standing annual cost dressed up as a one-off project.

Concentration. The Bank of England and the FCA found that a third of AI use cases in regulated firms are third-party implementations, up from 17% two years earlier, and that the top three model providers account for 44% of all named providers, against 18% in the previous survey. Concentration at that level is a pricing risk before it is a resilience risk. A buyer with one credible supplier is not negotiating; they are being informed.

Retention. Metered systems produce a transcript of everything, and in a regulated organisation that transcript is a record. Storage is trivial. Deciding what it is, how long it is kept, who can search it and what happens when a subject access request lands on it is not, and the cost of getting it wrong is not a storage cost.

How to build a forecast that survives contact with a rate card

Start from demand, not from licences. Count the interactions the process actually generates today, from the systems that already record them: contact volumes, ticket counts, document arrivals, form submissions. Those numbers exist and nobody has to estimate them.

Then map each interaction to the billable events it will produce. This is the step that is skipped, and it is where forecasts go wrong by an order of magnitude, because a single user question rarely costs a single unit. A grounded answer with a tool call and a follow-up is an answer, an action and a grounding event, not one of anything. Working through the rate card line by line for one representative interaction is an afternoon’s work and it is the most valuable afternoon in the whole project.

Then apply a range rather than a point. Model the case where adoption reaches half the target, the case where it reaches the target, and the case where it reaches three times the target because the tool turned out to be popular. Take the third seriously. Metered infrastructure is the only category of enterprise spending where the optimistic scenario is the expensive one, and a finance function that has never seen that shape before will assume the high case is the risk case rather than the success case.

Finally, write down what you will do at each threshold. Buy more capacity, throttle, restrict the agent to a narrower audience, or accept enforcement. Deciding that in advance is cheap. Deciding it during the month the meter runs out is not.

Why the published productivity case keeps outrunning the run cost

The ONS finding that ought to discipline every one of these business cases is not the adoption figure. It is what adoption has produced. Across all size bands, around half of businesses using AI report no change in their overall workforce headcount, and among the businesses using it specifically to improve operations, 63% report no change in worker headcount while 6% report a decrease and 1% an increase.

That is not an argument that the technology does not work. It is an argument about where the benefit lands. Capacity released inside a role does not become a saving unless somebody removes the role, reduces the establishment or stops buying agency cover, and none of those things happens because a licence was purchased. Meanwhile the meter runs from the first day.

So the honest form of the business case has two columns and one of them is empty at the start. Column one is the metered cost, which is knowable to the penny from the rate card and grows with success. Column two is the benefit, which is a claim about a decision the organisation has not yet made. A proposal that presents the second column as though it were as firm as the first is not being dishonest so much as unbalanced, and it is the commonest shape of paper reaching technology boards this year.

What to fix before the renewal

Instrument the meter before anyone else instruments the adoption. Consumption should be visible to the budget holder weekly, broken down by agent, so that a change in behaviour shows up while it is still a conversation rather than an invoice.

Write down which agents may be denied service and which may not, and buy capacity accordingly. Put the evaluation cost in the same business case as the licence, at the grade of the people who will actually do it. And read the rate card as a document that will be revised, because it has been revised, and the revision that matters will be the one that changes what counts as an action.

More of this analysis is indexed under AI and automation. The wider argument about ownership and run cost appears in what automation costs to run, and the question of whether to buy the runtime at all is examined in the build or buy decision.

The old joke about enterprise software was that the licence was the cheap part. It has stopped being a joke and become a pricing model.

Sources

  1. Microsoft, Copilot Studio published UK pricing microsoft.com
  2. Microsoft, Copilot Studio billing rates and management learn.microsoft.com
  3. Microsoft, licensing for agents powered by the standard harness learn.microsoft.com
  4. Bank of England and FCA, Artificial intelligence in UK financial services 2024, 21 November 2024 bankofengland.co.uk
  5. ONS, Artificial intelligence in UK businesses, 2023 to 2026 ons.gov.uk