London has the data. Citizens rarely see the benefit
TfL spends roughly £1m a year publishing open data and the assessed return is £90m to £130m a year. The city portal itself tells a much less flattering story.
Data and analytics
Most data failures in large organisations are not analytical failures. They happen at collection, in a configuration nobody owns, or in the gap between a documented method and the code that implements it.
The sharpest recent lesson about data quality in Britain was delivered by the organisation whose entire purpose is data quality. Sir Robert Devereux's independent review of the Office for National Statistics, published on 26 June 2025, traced three separate failures to their causes. A funding choice constrained the survey field force, which led directly to removing the pandemic-era boost to the Labour Force Survey. Problems in trade statistics stemmed from what the review calls known concerns about the efficacy of the underlying software configuration. Errors in the Producer Price Index reflected a mismatch between the methods that had been agreed and the coding which implemented them.
Not one of those is an analytics problem. One is a decision about fieldwork budgets, one is a system configuration that nobody owned, and one is the distance between a written method and its implementation. That is the shape of most data failure in any large organisation, and it is why this section spends more time on collection and lineage than on visualisation.
The review's recommendations are structural: a focused effort to improve core statistics, a change in how the organisation is led, and a fresh look at its governance. Note what is absent from that list. No new platform, no new tooling, no new dashboard.
The current enthusiasm for putting models on top of enterprise data has made the collection layer expensive to ignore. A model trained or prompted against records with inconsistent identifiers, undocumented deprecations and three competing definitions of a customer will produce confident answers that reproduce every one of those defects at speed and at scale.
British organisations tend to discover this in a specific order. The pilot works, because it ran against a curated extract that someone cleaned by hand. The rollout does not, because production data is not that extract. The two pieces that follow from that observation are the data infrastructure AI projects actually need and you cannot fine-tune your way out of bad records. The model coverage they connect to lives in AI and automation.
The Data (Use and Access) Act 2025 received Royal Assent on 19 June 2025. Its long title runs to most of a page and covers access to customer data and business data, services that verify facts about individuals, information standards, and more besides. Much of it takes effect through regulations made later rather than on the face of the Act.
That is precisely why the launch coverage was close to useless. The commercially significant details of a framework Act are settled in commencement regulations and in the schemes made under it, which arrive quietly, months apart, without a press notice. This desk reads those as they land and reports the ones that change what an organisation may do with data it holds. It does not report the Act again every time a consultancy publishes a summary of it.
There is an organisational tell for all of this that costs nothing to check. Find out who is accountable when a field changes meaning. In most large British organisations the answer is nobody, or it is an analytics team that consumes the field and cannot change the system that produces it. Lineage tooling does not fix that, because the problem is not that the change was invisible. It is that noticing it was not anyone's job.
Open data in Britain has a strong founding story and an unclear present. Portals exist and many are stale. The honest way to describe the state of it is to count: how many datasets on a given portal were updated in the past year, which publishers have stopped, and which have quietly changed licence. Counting is slower than asserting and it produces a number that can be rechecked next year by anyone. That is the approach taken in counting the open data portals still being updated, and the same discipline applies to everything published here. A figure without a route back to whoever produced it is decoration.
The founding argument for publishing public data was never mainly about transparency. It was that a dataset released once can be used by people the publisher will never meet, at a cost the publisher does not bear, which is an unusually good deal for a public body. Open data began well and then it stopped covers how that argument was won and where it lost momentum. London remains the most instructive case in the country, because it has published more than almost any comparable city and the benefit to residents is still difficult to point at, which is the subject of London has the data and citizens rarely see the benefit.
TfL spends roughly £1m a year publishing open data and the assessed return is £90m to £130m a year. The city portal itself tells a much less flattering story.
British law requires every public charge point to broadcast whether it is working, within 30 seconds. Almost nobody reads what that obligation produced.
The UK government's AI playbook sets out ten principles for civil servants. Not one of them is about whether the records being fed to the model are usable.
Companies House removed false or misleading information affecting 100,400 companies in a single year. That register is what British firms check clients against.
We queried data.gov.uk, sampled 6,000 of its datasets and then tested 400 of its download links. A fifth of them came back unusable and one in eight never connected.
The British open data programme was a delivery schedule with months attached, not a vision. Its central artefact has not been updated since March 2015.