Defence & Cyber AI
Militaries are fielding AI faster than they can assure it, and in the fifth domain the contest is already machine against machine. We advise on the discipline that spans both: traceability, test and evaluation, and responsible adoption, not targeting.
The sovereign stake
The window for shaping military AI norms is closing on an engineering timetable, not a diplomatic one. NATO's Maven Smart System is accredited on the Alliance's classified network; the CCW expert group must report before the 2026 Review Conference; national fielding programmes (the UK's Digital Targeting Web, the Pentagon's $13.4 billion AI and autonomy line) are already committed. Governments that lack sovereign capacity to test, evaluate and audit AI-enabled systems will be forced to accept vendors' and allies' claims on trust.
Assurance capacity is now a component of sovereignty. It takes years to build. The procurement decisions being signed this year assume it already exists. And the same gap has opened in the fifth domain: adversaries probe national networks in a state of persistent engagement, Washington's 2026 National Cybersecurity Strategy calls for agentic AI to scale network defence, and most states cannot yet test, bound or audit the autonomous defenders they are about to switch on.
The board-level stake
The dual-use boundary has dissolved, and many boards are on the wrong side of it without knowing. Frontier model providers now hold defence contracts; cloud, satellite, logistics and cyber firms already sit inside military supply chains; export-control regimes are tightening around AI compute and models. Chief AI officers should expect defence-grade assurance requirements (traceability, adversarial testing, human-accountability chains) to migrate into civilian regulation for critical sectors, as they did in aviation, nuclear and financial services.
Enterprises supplying dual-use AI carry obligations under emerging norms they have not yet priced. Knowing what your models do, evidencing it and bounding it is about to become a condition of doing business. For operators of critical infrastructure the question is sharper still: AI is joining your cyber defence either way, and when it acts at machine speed you will need evidence of what it did, and why, that survives the incident.
Military AI adoption has decisively outrun assurance. NATO bought Palantir's Maven Smart System in roughly six months, among the fastest procurements in Alliance history, and cleared it for the classified network in June 2026. The Pentagon's FY2026 request carries a first-ever dedicated line of $13.4 billion for AI and autonomy.
The UK is wiring a Digital Targeting Web for 2027. None of this is matched by equivalent investment in test and evaluation. This is the condition our founder coined 'Complex AI' for in 2019: a model carrying real consequence inside a multi-objective system, where the demo is the easy part. Exercises are run against a cooperative world. Electronic warfare, deception and distribution shift are not cooperative, and a tool that degrades quietly under them will be trusted longest at exactly the moment it is wrong.
Ukraine is not a story about autonomous weapons arriving. It is a story about engineering necessity setting norms before diplomats do. Terminal autonomy in strike and interceptor drones emerged because jamming severed the human link, not because anyone chose to delegate. Interceptor systems now down Shahed-type drones largely autonomously, by the hundreds. Both sides harvest battlefield data at a scale no Western test range can replicate. Two lessons carry. First, iteration speed is the capability; the platform is disposable. Second, whatever human-control policy a ministry writes will be arbitraged away at the front unless it is engineered in: logged, bounded and auditable by design. Retrofitting traceability after fielding has no record of success that we can find.
The governance architecture is three-tracked and thin. The CCW Group of Governmental Experts must report by the 2026 Review Conference; 156 states backed the UN General Assembly's autonomous-weapons resolution in First Committee in November 2025; the US-led Political Declaration holds at just over fifty endorsements; REAIM convenes but does not bind. All of it is declaratory.
None of it contains verification machinery, and the states building fastest are not waiting: the United States and Russia voted against the UNGA text. Our judgement: the workable near-term instrument is not a treaty but a common assurance baseline, Article 36-style legal reviews adapted for learning systems, plus mandated logging and audit trails. Traceability is the one requirement on which operators, lawyers and diplomats already converge. Build there.
The dual-use boundary has effectively dissolved. In 2025 the Pentagon's Chief Digital and AI Office placed awards worth up to $200 million each with Anthropic, OpenAI, Google and xAI; frontier commercial models now sit inside defence workflows on both sides of the Atlantic. That changes the assurance question for everyone. Governments must evaluate systems they did not build and cannot fully inspect.
Germany's hesitation over Maven, and its shortlisting of European alternatives, shows sovereignty anxiety hardening into procurement policy. Our advice is consistent: independence of evaluation matters more than provenance of code. A ministry that can rigorously test a foreign system is more sovereign than one that blindly trusts a domestic one. The Institute works on assurance, adoption discipline and governance. We do not advise on targeting.
Cyber is where the fusion of AI and security is already total. The fifth domain has become an algorithmic contest: offence automating, defence automating in response, and the line between cyber and conventional attack blurring as AI converges with electronic warfare. Washington has drawn the consequence in writing: the 2026 National Cybersecurity Strategy makes agentic network defence explicit doctrine, calling for AI to scale defence and disruption at machine speed, while the June 2026 Executive Order directs CISA to accelerate AI-enabled defence of critical infrastructure.
Our reading is that a machine-speed defender is itself a Complex AI deployment, an autonomous system acting inside the grid, the bank, the hospital, and it demands exactly the assurance discipline this frontier exists for: logged, bounded, auditable by design. The convergence is physical too. In March 2026, drones struck hyperscale data centres in the UAE and Bahrain, the first kinetic strikes on AI infrastructure. Defending compute, defending networks and defending the state are no longer separable problems. That is why this frontier treats them as one.
- What does test and evaluation mean for learning systems that keep changing after certification, and who re-certifies, how often, against what?
- How should Article 36 legal reviews be adapted for AI-enabled decision-support and continuously updated autonomy?
- What minimum standard of logging, traceability and audit should any military AI system meet before fielding, and can allies agree it as a common baseline?
- Where should governments draw the line between commercial frontier models and sovereign defence AI, and what national evaluation capacity does holding that line require?
- What assurance baseline should an agentic cyber defender meet before it is trusted inside critical national infrastructure?
- When attack and defence both run at machine speed, where does meaningful human control sit, and what evidence must survive the exchange?
Fellows for this frontier are being appointed.
People who have built, governed or operated real systems in this domain. The first five fellows are named, and the rest of the founding cohort follows in September.
Have you done this at scale in Defence & Cyber AI? We want to hear from you.
Put yourself forward