Frontier 1

Complex AI

A model is a component; a deployed AI system is a complex adaptive system, with the feedback loops, emergence and cascade risk that implies. Complex AI is the discipline of governing it: traceable, auditable, accountable across competing objectives.

1 · Why It Matters
For governments

The sovereign stake

The compliance calendar is moving; the system dynamics are not. From 2 August 2026 the European Commission can fine general-purpose AI providers up to 3% of global turnover, even as the Digital Omnibus defers high-risk obligations to December 2027. Governments are simultaneously the regulator and the largest deployer: the UAE is embedding a national AI system in federal decision-making; Saudi Arabia has declared 2026 its Year of AI. AI inside welfare, borders and justice is AI inside complex human systems, feedback loops included. The deferral buys time to build supervision infrastructure. States that treat it as a pause will meet the failure modes unprepared.

For enterprises

The board-level stake

Gartner expects more than 40% of agentic AI projects to be cancelled by end-2027, citing escalating costs, unclear value and inadequate risk controls. The pattern beneath: pilots are assessed as models, then deployed as systems. Once agents call tools, share memory and trigger other agents, failure stops being local: OWASP now names cascading agentic failure as a distinct security class. Boards should demand what auditors will demand within two years: decision provenance for every consequential automated action, tested rollback, and explicit ownership of multi-objective trade-offs. ISO/IEC 42001 and the pending NIST agent overlays sketch the direction. The firms building traceability before scale will be the ones still scaling in 2028.

2 · Our View

When our founder coined the term Complex AI in 2019, the argument was unfashionable: the hard problem is not model capability but what happens when a capable model is embedded in a live human system. That argument is now the operating reality. A deployed AI system is a complex adaptive system in the strict sense: components that learn, humans who adapt to the machine, and feedback loops in which the model reshapes the very behaviour it was trained to predict. Demo performance measures the component. Governance concerns the system. The gap between models that demo well and systems that govern well is, in our view, the defining engineering problem of this decade.

Complexity science stops being academic the moment you run agents in production. Capabilities and failures arrive nonlinearly: phase transitions, not gradients. Two agents, each behaving correctly against its local objective, can combine into a globally catastrophic outcome; security bodies now catalogue cascading agentic failure as a named class. You cannot unit-test emergence. Pre-deployment certification of components, the instinct inherited from safety-critical software, is necessary and insufficient. Assurance has to become continuous and system-level: instrumenting for drift, watching for attractor shifts, and rehearsing containment the way grid operators rehearse blackouts. Gartner projects that over 40% of agentic projects will be cancelled by end-2027, citing inadequate risk controls. We read that as the bill for skipping this step.

Entropy gives assurance its sharpest practical rule: an unlogged action is information destroyed, and no audit can recover it. Traceability is not paperwork; it is the only mechanism by which an institution can answer for a decision after the fact. Every consequential automated decision needs provenance: inputs, objective weights, tool calls, and who could have intervened. Regulators are converging on this instinct: the EU AI Act's record-keeping duties, NIST's draft control overlays for single- and multi-agent systems, ISO/IEC 42001's management-system discipline.

But standards trail deployment by roughly two years, and the omnibus deferral of high-risk obligations to December 2027 widens that gap. Institutions should build to the physics, not to the deadline. The UK's assurance market (£1.01 billion in 2024, £18.8 billion potential by 2035) shows evidence infrastructure becoming an industry in its own right.

What should actually be built? For governments: a live register of deployed systems, incident telemetry that crosses agency boundaries, staged deployment with tested rollback, and circuit breakers with named humans holding the switch. For enterprises: decision provenance as core infrastructure, budgeted like security, not bolted on like compliance. The Gulf will be an early proving ground: Saudi Arabia is consulting on an operational responsible-AI policy while building gigawatt-scale compute, and the UAE is wiring a national AI system into federal decision-making. States that industrialise deployment fastest will hit complex-system failure modes first, and will set the governing precedents. The institutions that endure will be those that treated AI as a system to be governed, not a model to be bought.

3 · Questions We Work On
  1. How do you detect a deployed AI system approaching a behavioural phase transition, before it crosses one?
  2. What is the minimum decision-provenance standard for an autonomous agent acting inside government or financial systems?
  3. When an agent's objectives conflict, who owns the trade-off, and how is that ownership evidenced to an auditor?
  4. What does continuous, system-level assurance look like when models update faster than certification cycles?
  5. Which circuit breakers contain multi-agent cascade without destroying the system's usefulness?
4 · In Numbers
From 2 August 2026 the European Commission can fine general-purpose AI providers up to €15m or 3% of global annual turnover
GPAI enforcement proceeds on schedule even as the EU's Digital Omnibus agreement (May 2026) provisionally defers Annex III high-risk obligations to 2 December 2027, pending formal adoption.
The UK AI assurance market comprises over 524 firms generating £1.01bn GVA (2024), with DSIT projecting up to £18.8bn by 2035
DSIT's Trusted Third-Party AI Assurance Roadmap (September 2025), backed by an £11m AI Assurance Innovation Fund opening Spring 2026.
Over 40% of agentic AI projects will be cancelled by end-2027 (Gartner, June 2025)
Cited causes are escalating costs, unclear business value and inadequate risk controls, even as Gartner expects 15% of day-to-day work decisions to be autonomous by 2028.
The Fellowship

Fellows for this frontier are being appointed.

People who have built, governed or operated real systems in this domain. The first five fellows are named, and the rest of the founding cohort follows in September.

Have you done this at scale in Complex AI? We want to hear from you.

Put yourself forward
Field-provenIndependentOn the record