AI and automation 8 min read
What enterprise AI agents actually shipped in 2026
Strip out the announcements and the pilots, and the enterprise agent estate is small, narrow and closely supervised. The public records show exactly how narrow.
The most useful public list of AI systems running in British organisations is not a market report. It is the government’s algorithmic transparency register, which now holds 141 published records, each one written by the organisation that owns the tool and each one carrying a phase: pre-deployment, private beta, public beta, production or retired.
Read it by phase rather than by headline and the enterprise agent story reorganises itself immediately. The Ministry of Justice published a record in August 2026 for an Office of the Public Guardian investigations assistant, described as AI-powered workflows supporting investigators with tasks such as financial transaction categorisation and analysis. Its phase is pre-deployment. The Department for Education’s proposed system for automating vacancy quality checks on apprenticeship adverts, published the same week, is also pre-deployment. The Department for Business and Trade’s export support chatbot is in private beta.
Against those, the Cabinet Office’s Assist, a generative tool for the government communications profession, is in production and has been since June 2025. So is the Information Commissioner’s Office chatbot for data protection fee enquiries, and so is the DVLA’s natural language voice routing on its contact centre.
That is the shape of it. What is in production is narrow, single-purpose and mostly conversational or classificatory. What is agentic in the sense the vendors mean, meaning software that plans, calls tools and acts, is sitting in pre-deployment with a named owner and a risk assessment attached.
The financial sector counted the same thing and got a number
Government is not a special case. The Bank of England and the FCA surveyed 118 regulated firms and published the results in November 2024. Of all AI use cases reported, 55% involved some degree of automated decision-making. Within that, 24% were semi-autonomous, which the report defines as able to make a range of decisions on their own while designed to involve human oversight for critical or ambiguous ones. Fully autonomous decision-making accounted for 2% of use cases. Automation with dynamic models accounted for another 2%.
Two per cent is the number to hold onto, because it is the only published, sector-wide, regulator-collected figure of its kind in Britain. Everything sold as an autonomous agent is competing for a slice of a category that a regulated industry had barely opened.
The same survey found 62% of use cases rated low materiality by the firms running them, and 16% rated high. High materiality clustered in general insurance, risk and compliance, and retail banking. That distribution is what you would expect from an industry putting its more capable systems where the money is and keeping them supervised, and it is inconsistent with the idea that agents are quietly running the back office.
Why the shipped things are shaped the way they are
There is a common structure to the systems that made it into production, and it is not a technical preference. It is a supervision preference.
Every production entry in the transparency register does one job. The ICO chatbot answers questions about data protection fees. It does not open cases, take payments or change a record. The DVLA voice system routes a call. Ofsted’s survey summarisation tool, in private beta, summarises free-text survey responses for inspectors, who then read them. In each instance the system’s output lands in front of a person before it does anything irreversible, which means the failure mode is a person reading something wrong rather than a wrong thing happening.
The National Cyber Security Centre made this explicit in guidance published on 15 May 2026, summarising joint international work on the careful adoption of agentic services. Its position is that organisations should start small and use agents only for low-risk tasks, and its list of what makes agents harder than earlier AI is short and precise: broader access to external systems and data, unpredictable behaviour where goals can be interpreted in unexpected ways, problems that are harder to spot, particularly when actions occur faster than humans can meaningfully review them, and behaviour that is difficult to explain after the fact.
The practical controls it recommends are the ones that any RPA programme in 2016 should have had and mostly did not. Grant the least privilege that works, for as short a period as possible. Bound the systems an agent may touch, the operations it may perform and the hours in which it may perform them. Do not issue credentials that outlive the task they were issued for. The NCSC also states the accountability position without hedging: a system may take an action, but humans remain accountable for the decision to deploy it, the access it was granted, the safeguards around it and the consequences of its operation.
An organisation that cannot name the person who holds those four things has not built an agent. It has built an incident with a delay on it.
What the pre-deployment records tell you
The pre-deployment entries are the most forward-looking public information available about enterprise agents in Britain, because an organisation only files one when it has decided to proceed and has something specific enough to describe.
Read them and a consistent scoping instinct appears. The Office of the Public Guardian’s investigations assistant is scoped to categorising financial transactions and supporting analysis for investigators, not to concluding investigations. The Department for Education’s vacancy checking proposal is scoped to review checks such as spelling and grammar, duplicate detection and missing or inconsistent content, not to accepting or rejecting adverts. In each case the agent produces material that a person then acts on, and the boundary is drawn at the point where an outcome becomes irreversible.
That is not timidity. It is the only design for which an accountable owner can currently be found. Ask a senior civil servant or a bank’s chief operating officer to put their name to a system that concludes cases without review and the answer is no, and it will remain no until somebody can show them an error rate they believe. Nobody has produced one, because producing one requires an evaluation regime that most organisations have not yet built and have not yet budgeted for.
The register also shows how long this takes. Cabinet Office Assist was published as a production system in June 2025. The entries filed in mid-2026 are mostly earlier in their life. A public organisation moving carefully takes eighteen months to get from a described intention to a supervised production service, and there is no reason to believe a large private organisation is faster, only that its progress is not published.
The vendors have already told you what they expect
If you want to know how much autonomy the software industry thinks it is about to sell, read the price list rather than the launch post.
Microsoft meters Copilot Studio in Copilot Credits, and its published rates are granular in a way that reveals the assumed workload. A classic answer, meaning a response an author wrote by hand, costs one credit. A generative answer costs two. An agent action, which covers triggers, deep reasoning and topic transitions, costs 5 Copilot Credits, and computer-using agents are billed at that same action rate. Agent flow actions, which are predefined sequences that run without reasoning at each step, cost 13 credits per hundred actions.
Look at the ratio rather than the digits. A hundred deterministic flow steps cost less than three reasoning actions. The vendor is pricing reasoning as the scarce thing and rote execution as almost free, which is a fair description of the underlying economics and also a strong hint about where it expects the volume to sit. The cheapest unit on the sheet is the one that does not think.
Salesforce arrives at a similar structure by a different route, metering Agentforce in credits where a standard agent action draws twenty of them and voice actions draw thirty. Its published UK pricing page quotes those credits in US dollars, which is worth noting for anyone building a sterling budget from it.
What is actually different this time
Three things, and none of them is capability in the abstract.
The first is that the integration problem got cheaper. A decade of automation failed on the seam between systems, and the current generation of tooling can read a schema, call an interface and handle a shape it was not explicitly programmed for. That is real, and it is the reason the pre-deployment records in the transparency register describe work that would not have been attempted in 2019.
The second is that the failure surface got wider in exactly the same motion. The thing that lets an agent adapt to an interface it has not seen is the thing that stops you being able to enumerate its behaviour in advance. Assurance for a rules-based automation is a test suite. Assurance for an agent is an evaluation regime with a sampling strategy, a held-out set and a review cadence, and somebody has to fund it every year rather than once.
The third is the one nobody puts in a business case. Adoption across British business has broadened without deepening. The Office for National Statistics reports that use of at least one AI technology among businesses with ten or more employees has risen to around 35% since late 2023, but the average adopting firm uses barely more technologies than it did at the start, and only 10% describe their use as extensive. Adoption varies enormously by sector, from 58% in information and communication to 13% in construction. A market where most adopters run one or two things is not a market that has agents managing processes end to end.
There is a measurement caveat worth carrying, because it cuts both ways. The ONS notes that estimates of adoption vary substantially with method, and that surveys of senior executives report far higher figures than surveys that ask about direct use. Bottom-up measurement captures what workers and specific functions actually do. Top-down measurement captures whether a firm uses any AI anywhere. A board that believes its organisation is an advanced adopter is often reading the second number while its operations are described by the first.
The question to ask a vendor
Not what the agent can do. Ask what it is permitted to do, by whom, and how that permission is revoked.
Then ask the harder version. When this system takes an action that turns out to be wrong at scale, which of your own staff finds out first, how long does that take, and what stops it in the meantime. The organisations with production entries in the transparency register can answer that, because filling in the record forces them to. The organisations with the most enthusiastic slide decks generally cannot.
The rest of the desk’s reporting on this sits under AI and automation. Arithmetic on the metering behind agent pricing appears in what these systems cost in their second year, and the longer argument about the operational discipline that carries forward from rules-based deployment is in our intelligent automation guide.
What shipped in 2026 is not the autonomous enterprise. It is a modest number of bounded tools with named owners, sitting in front of people who still make the decision. That is a less exciting sentence than the one on the conference stage, and it is the one the public records support.
Sources
- GOV.UK, Algorithmic Transparency Recording Standard, published records gov.uk
- Bank of England and FCA, Artificial intelligence in UK financial services 2024, 21 November 2024 bankofengland.co.uk
- NCSC, Thinking carefully before adopting agentic AI, 15 May 2026 ncsc.gov.uk
- Microsoft, Copilot Studio billing rates and management learn.microsoft.com
- ONS, Artificial intelligence in UK businesses, 2023 to 2026 ons.gov.uk