Data and analytics 6 min read
Open data began well. Then it stopped
The British open data programme was a delivery schedule with months attached, not a vision. Its central artefact has not been updated since March 2015.
The thing most often forgotten about the British open data programme is that it did not begin with a vision. It began with a delivery schedule.
On 31 May 2010 the Prime Minister wrote to every government department with a list and a set of months. Historic COINS spending data and the names, grades and salaries of senior civil servants earning over £150,000 in June. New central government ICT contracts in July. Central government spending over £25,000 in November. All central government contracts, Department for International Development projects over £500, local government spending over £500, local contracts and tenders, and crime data by location, in January 2011.
Whatever else can be said about that letter, it committed the government to specific artefacts on specific dates, and most of them appeared. Almost nothing in British data policy since has been written that way.
The argument was economic, and it was made honestly
The transparency framing has attached itself to this period retrospectively, helped along by the fact that the first release was about spending and salaries. It was not the main argument at the time and the documents say so.
The Cabinet Office’s Open Data White Paper, Cm 8353, published on 28 June 2012, set out three commitments: making it easier to access public data, making it easier for publishers to release data in standardised open formats, and “engraining a ‘presumption to publish’ unless there are clear, specific reasons (such as privacy or national security) not to do so”. The paper described open data as “an effective engine of economic growth, social wellbeing, political accountability and public service improvement”, in that order.
The reasoning underneath was sound and remains sound. A public body that publishes a dataset once is used by people it will never meet, for purposes it did not anticipate, at a cost it does not pay. There are very few arrangements in public administration with that shape. It does not require the publisher to be clever about what the data is for, which is fortunate, because publishers are reliably wrong about that.
The lasting invention was a licence, not a website
data.gov.uk gets the attention because it is the thing with a URL. The more consequential artefact was the Open Government Licence, developed by The National Archives, and it is worth reading because so few people who cite it have.
It grants a “worldwide, royalty-free, perpetual, non-exclusive licence to use the Information”. The permissions are to copy, publish, distribute and transmit; to adapt; and to “exploit the Information commercially and non-commercially”. The condition is attribution. The exclusions are narrow and specific: personal data, departmental logos, crests and the Royal Arms, military insignia, third party rights the provider cannot license, other intellectual property rights, and identity documents such as the British passport.
Two words carry most of the weight. Perpetual means the licence does not expire when a minister changes or a portal is retired, so a copy someone downloaded in 2013 is still lawfully theirs. Commercially means a business can build a product on it without negotiating, which removes the single largest reason a lawyer would tell a startup not to depend on public data.
A licence is cheap, it is portable, and it does not need a budget line to keep working. That is why the Open Government Licence has outlived almost every institution created alongside it.
The evidence base was thinner than the enthusiasm
By 2013 the programme had reached the stage every policy reaches, where somebody asks what it has been worth. The Shakespeare review of public sector information, led by Stephan Shakespeare and published on 15 May 2013 by the Department for Business, Innovation and Skills, made nine recommendations and rested on a market assessment produced by Deloitte.
That structure is the tell. Three years into a national programme, establishing what public sector information was worth still required commissioning a consultancy to estimate it, because the state had not instrumented the thing it was doing. Downloads were not consistently counted, reuse was not tracked, and no publisher was required to report either. A programme that asks other people to publish evidence about themselves and does not publish evidence about itself is in a weak position when the money gets tight.
The date it stopped is on a government web page
In October 2013 the Cabinet Office published the first iteration of the National Information Infrastructure. This was the serious version of the idea: not a portal of whatever departments happened to release, but a defined core of “the data held by government which is likely to have the broadest and most significant economic and social impact if made available and accessible outside of government”. Judgement of what belonged was to rest on likely benefit inside and outside government, with weight given to data used for emergency resilience, data published on a statutory basis, and data supporting an organisation’s public task.
Alongside it came an inventory of what government held but had not released. The page records that “over 3,900 unpublished datasets have been listed on the site so far”, with a mechanism inviting the public to say what releasing each would be worth.
That page was last updated on 24 March 2015. It has not been withdrawn, superseded or replaced. It simply sits there, the most ambitious statement of what British open data was supposed to become, eleven years past its last revision.
Why it stopped is not mysterious
There is a temptation to explain the stall by austerity, or by a change of government, or by the ordinary decay of enthusiasm. Those all contributed. None is the mechanism.
The mechanism is that nothing in the programme ever created a duty. A department that published spending data in 2011 and quietly stopped in 2016 broke no rule, failed no inspection and was reported to nobody. The white paper’s presumption to publish was a presumption, not an obligation, and it was enforced by ministerial attention. Ministerial attention is the least durable resource in Whitehall.
Compare that with the machinery built around statistics, or around company filings, or around the environmental information regime, all of which have a body whose job is to notice non-compliance. Open data was given a portal, a licence and a schedule of announcements, and no institution with a standing interest in whether the publishing continued.
The licence survived because it needed nobody. The schedule survived because it had dates. The rest depended on someone caring, and eventually nobody was paid to.
What that left behind, counted rather than asserted, is examined in counting the open data portals still being updated. The city that pushed hardest on this agenda, and where the distance between publishing and benefit can actually be measured, gets its own piece at beyond the smart city. Both sit under data and analytics.
Sources
- Prime Minister's letter to government departments on opening up data, 31 May 2010 gov.uk
- The National Archives, Open Government Licence version 3.0 nationalarchives.gov.uk
- Cabinet Office, Open Data White Paper: Unleashing the Potential, Cm 8353, 28 June 2012 gov.uk
- Department for Business, Innovation and Skills, Shakespeare review of public sector information, 15 May 2013 gov.uk
- Cabinet Office, National Information Infrastructure: first iteration, October 2013 gov.uk