Most AI strategies survive the boardroom and die in production. When we first published that sentence, it was an observation from the field. This brief is the evidence.

1. The claim

The past three years settled the question of capability. Models now write, reason, plan and act at a level that would have sounded absurd when I coined the term Complex AI in 2019. What the past three years did not settle is the question of operation. The models became general. The deployments did not.

We call the distance between the two the deployment gap: the space between a model that demos well and a system that governs well. It is measurable, it is wide, and it is not closing on its own. By RAND's estimate, more than 80 per cent of AI projects fail, twice the failure rate of IT projects that do not involve AI. That number should stop every minister and every board chair mid-sentence. AI is not merely hard. It fails at double the rate of the ordinary technology programmes that institutions already find hard.

This brief makes three claims and then does something about them.

First, failure is the base case, not the exception. The statistics below come from RAND, MIT, Gartner, BCG, McKinsey, S&P Global and the UK's National Audit Office. They disagree on magnitude. None of them disagrees on direction.

Second, the causes are institutional, not technical. When RAND interviewed 65 experienced data scientists and engineers about why their projects died, the root causes were misaligned purpose, inadequate data foundations, technology chosen before the problem, and organisations without the infrastructure to operate what they had bought. The model itself barely features. Our own experience of carrying systems into production says the same: most AI initiatives do not die because the model was weak. They die because the system around the model was never built.

Third, the fixes are known, unglamorous, and rarely funded. They are the disciplines of safety-critical engineering and institutional accountability: explicit trade-offs with named owners, external validation before scale, decision provenance, tested rollback, and a budget for the system around the model that matches the budget for the model. None of this is new. Almost none of it appears in the strategy decks we are shown.

What follows is the evidence, the failure patterns as builders see them, and a checklist designed to be carried into a cabinet meeting or a board pack.

2. The evidence

Start with the base rates, because they are worse than most decision-makers believe.

Gartner predicted in July 2024 that at least 30 per cent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs and unclear business value. By June 2025 it had extended the forecast to the next wave: more than 40 per cent of agentic AI projects cancelled by the end of 2027, blaming escalating costs, unclear business value and inadequate risk controls.

The realised numbers are keeping pace with the forecasts. S&P Global's 2025 enterprise survey found that the share of companies abandoning most of their AI initiatives jumped from 17 per cent to 42 per cent in a single year, with the average organisation scrapping 46 per cent of its proofs of concept before production.

MIT's NANDA initiative, after interviewing enterprise leaders and analysing 300 public deployments, concluded that 95 per cent of enterprise generative AI pilots deliver no measurable profit-and-loss impact, set against what the underlying report puts at 30 to 40 billion dollars of enterprise investment. BCG's global survey of 1,000 senior executives found that only 26 per cent of companies have built the capabilities to move beyond proof of concept, and only 4 per cent consistently generate significant value. McKinsey's 2025 global survey completes the picture: 88 per cent of organisations now use AI somewhere, yet only 39 per cent can attribute any enterprise-level earnings impact to it at all.

The most recent reading comes from the people running the systems. In April 2026 Gartner published a survey of 782 infrastructure and operations leaders, which found that only 28 per cent of AI use cases fully succeed and meet their return expectations while 20 per cent fail outright, and that among leaders reporting at least one failure, 57 per cent put it down to expecting too much, too fast. That is the deployment gap described by the people standing in it.

Government is not the exception. It is the same pattern with higher stakes. The UK National Audit Office's March 2024 survey of 87 public bodies found that 37 per cent had deployed any AI, 70 per cent were piloting or planning, and only 21 per cent had an AI strategy. The gap between announcement and operation is, if anything, wider in the public sector, because withdrawals are quieter than launches.

Statistics establish the pattern. Cases establish the mechanism. Here are four, each documented, each expensive, and each dead for reasons that had almost nothing to do with model quality.

Case one: sixty-two million dollars, no system

MD Anderson Cancer Center and IBM spent four years building the Oncology Expert Advisor, a Watson-powered clinical decision tool. A University of Texas System audit found that the original 2.4 million dollar, six-month contract was extended twelve times, reaching 39.2 million dollars in fees to IBM alone, with total project costs of at least 62 million dollars, procured outside standard processes.

The tool was benched in 2016 without ever being used on a patient. The proximate cause of death is the detail every builder should memorise: the system had been developed against the hospital's old medical records platform and was never integrated with the new one after the institution changed systems mid-project. The model was arguably the most famous AI in the world at the time. It died of an integration dependency. IEEE Spectrum's post-mortem of the wider Watson Health programme found the same anatomy repeated across deployments.

Case two: the algorithm that helped bring down a government

For years the Dutch tax administration used an algorithmic risk-classification model to flag childcare benefit claims for fraud investigation. The Dutch data protection authority found that the tax administration had unlawfully held data on applicants' dual nationality and used nationality as an indicator in its risk model, and fined it a record 2.75 million euros.

Families were wrongly branded fraudsters and forced to repay tens of thousands of euros; a parliamentary inquiry titled Unprecedented Injustice concluded that fundamental principles of the rule of law had been violated, and the government earmarked 500 million euros in compensation for more than 20,000 parents. On 15 January 2021, the entire Dutch cabinet resigned. Note what failed. The model did exactly what it was built to do: maximise fraud detection. Nobody with authority had written down what it was allowed to trade away to do it, and no institution could reconstruct or defend the individual decisions when challenged. This is what a single-objective system does inside a multi-objective institution.

Case three: scrapped rather than defended

For years the UK Home Office streamed visa applications through an algorithm that assigned red, amber or green risk ratings, using nationality as an input. When the Joint Council for the Welfare of Immigrants and the legal campaign group Foxglove brought a judicial review arguing the tool was discriminatory under the Equality Act 2010, the Home Office scrapped the system from 7 August 2020, before the case was ever heard.

It promised a redesign by the end of October 2020, with interim decisions made on person-centric attributes and nationality excluded altogether. The operational lesson sits in the sequence of events: the government chose abandonment over disclosure. A system whose logic cannot be explained in court will not be defended in court. I wrote in the founding essay that systems which can explain their decisions get defended by their institutions, and that black boxes get quietly switched off. This case is that sentence with a date on it.

Case four: scaled before it was validated

The Epic sepsis model shipped inside the most widely used electronic health record in the United States and was adopted by hundreds of hospitals. When University of Michigan researchers ran a major independent validation across 27,697 patients, they found the model caught only 33 per cent of sepsis cases, missing two thirds, while firing alerts on 18 per cent of all hospitalisations. The results were published in JAMA Internal Medicine in 2021, years after the model had reached national scale. Hospitals had bought a vendor benchmark, not a validated system. Every clinician who learned to ignore those alerts was the deployment gap operating at the bedside.

One coda, because it removes the last comfortable excuse. Amsterdam spent years building Smart Check, a welfare-fraud model designed to be everything the Dutch national scandal was not: an explainable model, an entry in the city's algorithm register, extensive bias testing, academic oversight, community consultation.

In the pilot it still could not be made both fair and effective, and the city shelved it. Lighthouse Reports' joint investigation with MIT Technology Review records that the city's own Participation Council had said the experiment touched citizens' fundamental rights and should be discontinued, advice the city overrode. Doing everything right inside the project cannot rescue a deployment that skipped the question of whether. That question belongs to ministers and boards, and to no one else.

One more, because 2026 supplied a case the others could not. In April South Africa withdrew its draft National AI Policy after roughly a tenth of the academic sources in its reference list were found not to exist; the minister’s explanation was that AI-generated citations had been included without verification.

Two officials were suspended for failing to declare that they had used AI, and two more were suspended at Home Affairs over fabricated sources in a separate white paper. The document meant to govern the state’s use of AI was undone by ungoverned use of it. Every question in the next section applies to a policy process as surely as to a production system.

3. Why deployments fail: five patterns from the field

We have built, shipped and operated these systems. The failure modes are not mysterious. They recur so reliably that we can name them.

Pattern one: the pilot was never the product. A pilot succeeds on clean data, with a narrow objective, hand-held by its builders and graded by people who want it to work. Production is the opposite of every one of those conditions: legacy integrations, adversarial users, data that drifts, and nobody in the room who remembers the demo. The Watson project died of an unbudgeted integration with a records system. The honest framing is that a pilot proves the model and proves nothing about the system, and the system is most of the work.

MIT's finding that 60 per cent of organisations evaluated enterprise AI tools, 20 per cent piloted, and 5 per cent reached production is what a funnel looks like when everyone budgets for the first stage.

Pattern two: one objective in, many objectives out. Real institutions optimise for several things at once: cost against care, speed against fairness, detection against due process. A model optimises for whatever it was given. Deploy a single-objective model into a multi-objective institution and the model will silently win, until the day it catastrophically loses. The Dutch system maximised fraud detection and traded away the rule of law, and a government fell. The failure was not that the trade-off existed. Trade-offs always exist. The failure was that no named human had accepted it in writing before the system went live.

Pattern three: the orphaned system. Models are launched by project teams and killed by operations. The consultants leave, the sponsor is promoted, the retraining budget is cut in year two, and the system decays until an incident review discovers that nobody owns it. RAND's interviewees put inadequate infrastructure to manage data and deploy finished models among the leading root causes of AI project death, and the National Audit Office found UK departments piloting enthusiastically while fewer than a quarter had a strategy for what they were piloting.

A model is a purchase. A system is a commitment with a headcount, a maintenance budget and a named owner whose performance review depends on it. If those three things are missing, the deployment has already failed; the calendar just has not caught up.

Pattern four: validation stopped at the vendor's benchmark. The Epic sepsis model reached hundreds of hospitals before anyone independently checked it against reality, and reality returned a sensitivity of 33 per cent. This pattern is endemic because incentives point the same way on both sides of the transaction: vendors report the benchmark that sells, and buyers lack the capability to test on their own data, so procurement substitutes brand for evidence.

There is a simple discipline here, and it is not negotiable in anything I would put my name to: no consequential AI system goes live on a population it has not been validated against, and validation is performed or audited by someone who does not profit from a positive result.

Pattern five: accountability was designed in last. Auditability is treated as a compliance layer to be added before the regulator visits. It is actually the survival layer. The Home Office abandoned its visa tool rather than defend its logic in open court. Air Canada argued before a tribunal, in effect, that its own customer-facing chatbot was "a separate legal entity that is responsible for its own actions", and lost, because the argument was absurd: the deploying institution answers for the system, always.

An unlogged decision is information destroyed, and no audit can recover it. Systems built with decision provenance from day one (inputs, objective weights, the action taken, who could have intervened) get defended by their institutions when challenged. Everything else gets switched off, and the switch-off is booked as an AI failure when it was in fact an accountability failure that AI merely exposed.

4. What to do: ten questions before the money moves

This is the section to photograph. These are the ten questions I ask before signing, and they are for a minister approving a national deployment or a board approving an AI budget. Every dead system in this brief would have been caught by at least one of them.

  • 1. Who owns this system after the launch party? A named individual, in post for the life of the system, with budget and authority. If the answer is a programme team that dissolves at go-live, do not sign.
  • 2. What are we trading away, and who accepted the trade? Every objective the system optimises implies objectives it sacrifices. Demand the list in writing, with a signature against each trade-off. The Dutch cabinet resigned over an unsigned one.
  • 3. Where is the budget for the system around the model? Data pipelines, integration, escalation paths, human oversight, retraining cadence. If the budget is mostly model licences and the integration line is thin, you are buying a 62 million dollar pilot.
  • 4. Has it been validated on our population, by someone who does not profit from a yes? A vendor benchmark is marketing. Independent validation on your own data, before scale, is the minimum.
  • 5. Can we reconstruct any consequential decision this system makes? Inputs, weights, action, and the human who could have intervened. If the answer is no, you are signing for a system no lawyer can defend and no audit can reconstruct.
  • 6. Have we rehearsed turning it off? A tested rollback to a working manual or legacy process, exercised the way grid operators rehearse blackouts. An off-switch that has never been pulled is a hypothesis.
  • 7. What are the kill criteria, and who holds the switch? Written before launch: the measurable conditions under which the system is suspended, and the named human with authority to act on them without a committee.
  • 8. Who will stand in public and defend this system's decisions? Before a court, a select committee, a bereaved family. If no official or executive will own that podium, the institution has already told you it does not trust the system.
  • 9. Are the people funded? Trained users, staffed escalation, clinicians or caseworkers who will not learn to ignore the alerts. A system that its users route around has failed regardless of what the dashboard says.
  • 10. Should this exist at all? Not every process improves with prediction attached. Amsterdam's own citizen council said stop; the city spent years learning it was right. The whether question is the one only you can answer, and it comes first, not last.

Ten yeses do not guarantee a deployment. Any single no, unresolved, guarantees the alternative.

5. What the Institute will measure next

The deepest problem in this field is not the failure rate. It is that the failure rate has to be assembled from surveys, court filings and audit reports, because nobody systematically measures deployment outcomes. Launches are press releases; withdrawals are silence. The Institute's founding research programme is built to correct that, and we commit to four instruments.

A deployment survival index. We will track cohorts of publicly announced government and major enterprise AI deployments, across the UK, the EU, the US and the Gulf, and report what fraction remain in production at 12, 24 and 36 months, with documented cause of death for those that do not. The industry publishes adoption curves. We will publish survival curves.

The system-to-model spend ratio. Across deployments that survive, what was actually spent on integration, oversight, data infrastructure and operations for every unit spent on the model itself? Our field experience says the ratio is far higher than any budget we are shown. We intend to put a number on it and give ministers and boards an evidence-based planning multiple.

A minimum decision-provenance standard. Regulation is converging on record-keeping duties without specifying what a reconstructable decision requires in practice. We will publish and defend a floor: the smallest set of logged facts (inputs, objective weights, action, intervention authority) that lets an institution answer for an automated decision after the fact, in domains where decisions touch citizens.

A register of documented withdrawals. The cases in this brief are known because a journalist, a tribunal or an auditor happened to look. We will maintain and publish a running register of documented public-sector AI withdrawals and their causes, because the field cannot learn from failures it is not allowed to see.

The deployment gap is not a passing condition of an immature market. It is what happens, reliably, when capable models meet complex institutions without the engineering and accountability disciplines that every other safety-critical field learned the hard way. The states and enterprises that close the gap will compound an advantage for decades. The ones that keep funding pilots and calling them strategies will keep appearing in briefs like this one.