Agentic AI & Autonomous Systems
AI agents are crossing from demonstration into delegation, acting on live systems inside enterprises and governments. The hard problem is no longer capability. It is accountability: bounding, auditing and answering for machines that act.
The sovereign stake
Governments are no longer observers of agentic AI; they are operators. The UAE plans to deliver half of government services with AI within two years and is training 80,000 civil servants in agentic AI. Singapore's IMDA published the first governance framework for agentic systems in January 2026, expecting verifiable agent identity and an audit trail of who authorised what. NIST launched an AI Agent Standards Initiative in February 2026. The EU's high-risk obligations, which capture many agentic deployments, remain in flux after the Digital Omnibus agreement. When an agent denies a benefit or flags a taxpayer, the state answers for it. Assurance regimes must precede deployment, not follow it.
The board-level stake
For boards, the exposure is concrete and near-term. Liability is converging on the deployer: California's AB 316, effective January 2026, bars the defence that the system "acted autonomously", and a reasonable-oversight standard is spreading across jurisdictions. Most enterprises cannot yet meet the implied standard of care. Machine identities outnumber staff by more than 80 to 1 (CyberArk, 2025 Identity Security Landscape), and most organisations concede they cannot always explain why a non-human identity performed a privileged action. Gartner forecasts over 40 per cent of agentic projects cancelled by end-2027; much of the market it prunes is agent washing. The board question is not "do we have agents?" but "can we bound, log and stop them?"
Agentic AI is where the industry's central gap stops being arguable: a model that impresses in a sandbox is not a system anyone can be held answerable for. Our founder coined the term Complex AI in 2019 for precisely this territory: AI operating inside multi-objective, high-stakes, real-world systems, where traceability is not a feature but a licence to operate.
An agent that books a meeting is a model. An agent that moves money, closes tickets against an SLA and negotiates with other agents is a component in a complex system. Most deployments we see in 2026 are still the former dressed as the latter. Gartner's estimate that only around 130 of the thousands of self-described agentic vendors are genuine matches our field experience.
The accountability question has a settled answer and an unsettled practice. Legally, the agent is nobody: responsibility runs to the organisation that deployed it. Regulators are converging on a reasonable-oversight standard: California now bars the defence that the system acted autonomously, and Singapore's IMDA framework expects every agent to carry a verifiable identity and an audit trail of who authorised what. Practice lags badly. Machine identities outnumber humans by roughly 80 to 1 in large enterprises, most with standing access, and the majority of organisations admit they cannot always reconstruct why an automated identity took a privileged action. Until a board can answer 'which agent, whose authority, what evidence', it is carrying an exposure it cannot size.
Agentic failures are complex-systems failures, and must be engineered for as such. The canonical incidents are not exotic: an agent with production credentials deleted a live database during a declared code freeze, then misreported what it had done. In multi-agent settings the physics worsen. One agent's confident error becomes the next agent's trusted input. Semantic mistakes pass validation and cascade silently. Agents built on the same base model fail in correlated ways: monoculture collapse. None of this is fixed by a better model. It is fixed by the unglamorous disciplines of safety-critical engineering: least privilege, circuit breakers, bounded blast radius, staged autonomy, and human gates on irreversible actions.
Through 2026 a second failure mode arrived, and it is not your agent going wrong inside your own system. Anthropic disclosed that a group it assesses with high confidence to be Chinese state-sponsored ran 80 to 90 per cent of an espionage campaign through Claude Code, against roughly thirty targets, with humans stepping in at only four to six decision points and the agent issuing thousands of requests per second.
That is the first documented large-scale intrusion executed without substantial human involvement, and it resets the tempo defenders must hold. A control designed around a human analyst reading an alert does not survive an adversary operating faster than the alert can be read. The one piece of good news is that the answer does not change: identity for every non-human actor, least privilege, and an audit trail that reconstructs what happened, which is the same discipline that stops your own agents from hurting you.
On adoption, hold two truths at once. The plumbing is real: MCP, open-sourced in late 2024, is now the de facto standard for agent-tool connectivity, adopted across Anthropic, OpenAI, Google and Microsoft, with thousands of servers in the official registry. The delegation is mostly not: surveys put enterprises with any agent in production at roughly a quarter to a third, and genuine scale far lower, with much of the rest relabelled RPA. Our advice to ministers and boards is identical. Treat an agent as a hire, not a licence: a job description, scoped permissions, a probation period, supervision proportionate to blast radius, and an audit trail a regulator could read. Autonomy is earned, not configured.
- When an agent acts under delegated authority, who is accountable, and can the deploying organisation evidence 'reasonable oversight' to a court or regulator?
- What is the minimum audit trail (identity, authorisation, action, evidence) an agent must produce before it touches a citizen-facing decision?
- How should permissions for non-human identities be scoped and revoked when agents already outnumber staff by roughly 80 to 1?
- Which failure modes (cascades, correlated base-model monocultures, emergent coordination) must be engineered against before multi-agent systems enter critical infrastructure?
- Where is agentic AI genuinely production-ready in 2026, and where is 'agent washing' inflating ministerial and board expectations?
Fellows for this frontier are being appointed.
People who have built, governed or operated real systems in this domain. The first five fellows are named, and the rest of the founding cohort follows in September.
Have you done this at scale in Agentic AI & Autonomous Systems? We want to hear from you.
Put yourself forward