Foundry4

Data and analytics

Data and analytics

Most data failures in large organisations are not analytical failures. They happen at collection, in a configuration nobody owns, or in the gap between a documented method and the code that implements it.

The sharpest recent lesson about data quality in Britain was delivered by the organisation whose entire purpose is data quality. Sir Robert Devereux's independent review of the Office for National Statistics, published on 26 June 2025, traced three separate failures to their causes. A funding choice constrained the survey field force, which led directly to removing the pandemic-era boost to the Labour Force Survey. Problems in trade statistics stemmed from what the review calls known concerns about the efficacy of the underlying software configuration. Errors in the Producer Price Index reflected a mismatch between the methods that had been agreed and the coding which implemented them.

Not one of those is an analytics problem. One is a decision about fieldwork budgets, one is a system configuration that nobody owned, and one is the distance between a written method and its implementation. That is the shape of most data failure in any large organisation, and it is why this section spends more time on collection and lineage than on visualisation.

The review's recommendations are structural: a focused effort to improve core statistics, a change in how the organisation is led, and a fresh look at its governance. Note what is absent from that list. No new platform, no new tooling, no new dashboard.

Everything downstream inherits this

The current enthusiasm for putting models on top of enterprise data has made the collection layer expensive to ignore. A model trained or prompted against records with inconsistent identifiers, undocumented deprecations and three competing definitions of a customer will produce confident answers that reproduce every one of those defects at speed and at scale.

British organisations tend to discover this in a specific order. The pilot works, because it ran against a curated extract that someone cleaned by hand. The rollout does not, because production data is not that extract. The two pieces that follow from that observation are the data infrastructure AI projects actually need and you cannot fine-tune your way out of bad records. The model coverage they connect to lives in AI and automation.

The legal ground moved in June 2025

The Data (Use and Access) Act 2025 received Royal Assent on 19 June 2025. Its long title runs to most of a page and covers access to customer data and business data, services that verify facts about individuals, information standards, and more besides. Much of it takes effect through regulations made later rather than on the face of the Act.

That is precisely why the launch coverage was close to useless. The commercially significant details of a framework Act are settled in commencement regulations and in the schemes made under it, which arrive quietly, months apart, without a press notice. This desk reads those as they land and reports the ones that change what an organisation may do with data it holds. It does not report the Act again every time a consultancy publishes a summary of it.

There is an organisational tell for all of this that costs nothing to check. Find out who is accountable when a field changes meaning. In most large British organisations the answer is nobody, or it is an analytics team that consumes the field and cannot change the system that produces it. Lineage tooling does not fix that, because the problem is not that the change was invisible. It is that noticing it was not anyone's job.

Open data, honestly counted

Open data in Britain has a strong founding story and an unclear present. Portals exist and many are stale. The honest way to describe the state of it is to count: how many datasets on a given portal were updated in the past year, which publishers have stopped, and which have quietly changed licence. Counting is slower than asserting and it produces a number that can be rechecked next year by anyone. That is the approach taken in counting the open data portals still being updated, and the same discipline applies to everything published here. A figure without a route back to whoever produced it is decoration.

The founding argument for publishing public data was never mainly about transparency. It was that a dataset released once can be used by people the publisher will never meet, at a cost the publisher does not bear, which is an unusually good deal for a public body. Open data began well and then it stopped covers how that argument was won and where it lost momentum. London remains the most instructive case in the country, because it has published more than almost any comparable city and the benefit to residents is still difficult to point at, which is the subject of London has the data and citizens rarely see the benefit.

Latest in Data and analytics

Open data began well. Then it stopped

The British open data programme was a delivery schedule with months attached, not a vision. Its central artefact has not been updated since March 2015.