Foundry4

Cloud and infrastructure 6 min read

There is no first time. There is only the rehearsal

The 2019 census rehearsal found blocks of flats with no flat numbers in the address frame. TSB ran a performance test in one data centre, and the FCA priced that.

On 13 October 2019 the Office for National Statistics ran a census. Not the census. A census, in four local authority areas, seventeen months before the real one, with a real field force, real suppliers, real letters through real letterboxes and no statutory purpose whatsoever.

It covered 331,359 households across Carlisle, Ceredigion, Hackney and Tower Hamlets, and produced 101,774 returns, a rate of 30.7%, of which 83,316 came in online. Nobody was counted. Everything was learned.

That is the shape of competent delivery in a genuinely complex domain, and it is almost the opposite of what the phrase getting it right first time is normally taken to mean. Getting it right first time describes an outcome. The only method that reliably produces it is doing the thing repeatedly beforehand, in conditions close enough to the real ones that the differences can be written down.

What a rehearsal finds that a review cannot

The census address frame is built on AddressBase Premium, the Ordnance Survey and GeoPlace product, and it underpins the entire operation. Every invitation letter, every questionnaire tracking record and every follow-up visit depends on it.

Inspected in a database, the frame looked fine. Put through a postal run, it produced this: of the 331,359 letters addressed and posted, 3,473, about 1%, came back undelivered as addressed, and some addresses for blocks of flats had no flat numbers, which the evaluation records as making them difficult to follow up in the field.

Consider what kind of defect that is. It is not a data error in the sense of a wrong value. Every one of those records is a legitimate address, correctly geocoded, in a dataset maintained to a national standard. The fault only exists in the relationship between the record and a person standing in a communal entrance hall trying to work out which of eleven doors to knock on. No schema validation finds it. No sample query finds it. A rehearsal finds it, because a rehearsal is the only thing that puts a human being in that hallway.

The return rates make the second point. They ranged from 52% in Ceredigion and 48% in Carlisle to 25% in Tower Hamlets and 22% in Hackney. Any competent designer could have predicted that dense urban areas would be harder. Nobody could have predicted the size of the gap, and the size is what determines how many field officers you hire, where you put them and how long the operation runs. ONS also recorded, honestly, that the October timing meant field work happened on dark nights rather than spring evenings, so the rates cannot be read across to the real census. A rehearsal that states its own divergences from the real thing is worth more than one that pretends to have none.

The organisation’s own list of resultant actions is the tell that it worked. Train the field force to engage rather than simply to knock on doors. Develop the management information so that problems surface earlier. Increase the clerical checks on the address frame, and improve it using Census Coverage Survey data, administrative data and checks against Google Maps and Street View. Not one of those is a change to the collection software. All of them are things you can only discover by watching people fail at the actual job.

The counter-example has a price attached

Britain’s most expensive recent lesson in rehearsal fidelity was delivered in retail banking, and the regulator’s account of it is unusually specific about the moment the decision went wrong.

TSB migrated its customers onto a new banking platform in April 2018. The programme had planned for what it called dress rehearsals, which the Financial Conduct Authority’s final notice of 20 December 2022 defines as migration acceptance cycles carried out in real time, to confirm that the combination of software, people and processes could achieve the migration inside the window available. That is the correct definition of a rehearsal and TSB wrote it down. One of the programme’s guiding principles required a clean migration acceptance cycle before dress rehearsals began.

The platform ran across two data centres, intended to operate in what the notice calls Active-Active configuration, where both sites carry customer sessions simultaneously and either can absorb the whole load if the other fails. TSB had no pre-production environment capable of stressing the platform, so the plan was to run end-to-end performance testing in both data centres, in the production environment, in the configuration the bank would actually be running.

At the Migration Testing Forum on 28 February 2018, a proposal was made to test using only one data centre, so that services already migrated, including cash machines, would stay available. TSB initially rejected it, on the grounds that the test should reproduce the conditions expected at migration. It then accepted a counter-proposal to do exactly that, reasoning that cash machines would otherwise be unavailable for three or four hours, that the two data centres had been assured to be identically configured, and that passing a load test on half the capacity implied comfort about the full amount.

The two data centres were not identically configured. The global load balancer and the session security component were both wrong, sessions failed to persist between sites, and following migration there were periods when digital channels were entirely unavailable to customers.

The FCA’s finding is worth quoting for its precision. The environment in which testing was conducted before go-live should have simulated the post-migration environment as closely as possible, and instead a decision was taken that made the testing and production environments distinctly different. Had the performance testing been conducted in Active-Active configuration, the notice states, it is likely the issue would have been identified and the consequences for customers avoided.

TSB received 225,492 complaints and paid £32,705,762 in redress. The FCA imposed a penalty of £29.75m, reduced from £42.5m for settlement, with the Prudential Regulation Authority adding its own.

The decision was not a technical decision

Here is the part that transfers to every complex programme regardless of sector.

The FCA records that the decision to abandon Active-Active testing was taken informally, outside TSB’s governance structure, was not documented, and was not escalated. The bank did not believe it represented a risk and considered it a matter of purely technical judgement. Because it was never identified as a risk, no mitigations were considered for it.

That is how rehearsal fidelity is almost always lost. Not by a board resolving to skip the rehearsal, which no board would do, but by a series of individually reasonable technical accommodations, each made by someone qualified to make it, each shaving the distance between the rehearsal and the real thing, none of them recorded anywhere that a risk committee would look. By the time the exercise runs, it is testing a system that will not exist on the day.

The defence is procedural and it costs almost nothing. Write down, at the start, what the rehearsal is meant to reproduce. Maintain a register of every accepted divergence between the rehearsal environment and the production one, with a named owner for each. Put scope reduction inside the governance structure rather than inside the engineering conversation, so that removing a condition from the rehearsal is a decision with a signature on it. And insist that the rehearsal runs against the clock, because a migration that succeeds in eighteen hours when the window is twelve has failed.

Why this is the only honest answer to complexity

The instinct in a genuinely complex domain, whether that is a national statistical operation, a spatial data programme with hundreds of contributing authorities, or a core banking replacement, is to respond to complexity with more analysis. Analysis is necessary and it has a ceiling. What it cannot do is tell you what happens when the address frame meets a residential block, or when four thousand field officers who have never used the device try to use it in the dark, or when two data centres that the documentation says are identical turn out not to be.

Those facts are not discoverable by thinking harder. They are only discoverable by doing the thing, at some scale, before it counts, and by treating what you find as the actual product of the exercise rather than as an inconvenience to be explained away.

ONS budgeted seventeen months and four local authorities for that knowledge and got a census that worked. TSB saved three or four hours of cash machine availability.

Why the records underneath programmes like these are so often defective before anyone tests anything is covered at you cannot fine-tune your way out of bad records. The same failure at organisational rather than technical scale appears in why transformation stalls at the second department. This desk’s wider work sits under cloud and infrastructure.

Sources

  1. Office for National Statistics, 2019 collection rehearsal evaluation report for Census 2021, England and Wales ons.gov.uk
  2. Financial Conduct Authority, final notice to TSB Bank plc, 20 December 2022 fca.org.uk