Edge AI & Compute
Intelligence is moving to the device while the compute that feeds it collides with the physics of the grid. We work both ends of the stack: what to run at the edge, what to build behind the meter, and what a sovereign token actually costs.
The sovereign stake
An AI strategy that has not been costed against the grid is a press release. Grid queues, transformer lead times and water, not model quality, set the pace of national compute programmes, and the IEA expects global data-centre electricity demand to roughly double by 2030. Meanwhile the sensitive half of government AI, including defence, border systems, health and critical infrastructure, increasingly cannot leave the jurisdiction, or the device. Edge inference gives ministers a third option between renting foreign cloud and building gigawatt campuses: sovereignty per workload. The governments moving fastest, notably in the Gulf, treat compute siting, grid planning and AI procurement as one portfolio. Most others still run them as three.
The board-level stake
Inference is now a line item boards can see, and it compounds with usage in a way training never did. Frontier-class token prices have fallen by orders of magnitude since 2022, yet total spend keeps rising as agentic systems multiply calls. The practical response is architectural: route the bulk of routine work to small, task-tuned models, whether on NPUs the business already owns or on dedicated inference capacity, and reserve frontier models for decisions that warrant them. Done well, this cuts cost, latency and data exposure simultaneously. It also changes governance: a fleet of small models demands versioning, evaluation and audit discipline most enterprises have not yet built.
The centre of gravity in AI has moved from training to inference, and inference has geography. By most industry estimates inference now accounts for roughly two-thirds of all AI compute, up from a third in 2023. Where a token is generated determines its latency, its legal jurisdiction and its failure modes. A well-tuned 3 to 8 billion-parameter model on a modern NPU beats a frontier model across a contested or expensive link for most operational decisions. The hard problem is not the model; it is the routing layer that decides what is handled locally and what escalates to the cloud. That layer is where governance actually lives, and few institutions yet engineer it deliberately.
Just under 3 per cent of world electricity by 2030, and slightly more than Japan consumes in total today. The constraint is not the projection but the connection: high-voltage transformer lead times now run to five years, and capacity queues are measured in gigawatts that have no delivery date.
Power, not silicon, is the binding constraint on the centralised stack. The IEA expects global data-centre electricity to roughly double to around 945 TWh by 2030. ERCOT's large-load queue nearly quadrupled to 226 GW in a single year. High-voltage transformer lead times now run to five years, and analyst tallies put 2026 hyperscaler capex between $600 and $690 billion, much of it chasing grid connections that do not yet exist.
The strategic consequence is underpriced: every workload pushed to a device is demand permanently removed from the queue. Edge deployment is a demand-side instrument of energy policy. We have yet to meet an energy ministry that treats it as one. The first government to do so will have connected its AI strategy to its grid strategy; most still run them from separate buildings.
The Gulf is executing the most decisive buildout outside the United States. Stargate UAE's first 200 MW is on track for the third quarter of 2026, inside a planned 5 GW Abu Dhabi campus; Saudi Arabia's HUMAIN is scaling towards multi-gigawatt capacity on the back of $23 billion in vendor partnerships. The region's advantage is real: energy-first siting, single-decision-maker speed, capital depth. But say it plainly: a gigawatt campus running closed weights under another state's export licence is hosting, not sovereignty. Sovereignty is the full stack: weights, orchestration, audit and people. The edge completes it. A nation that pairs sovereign campuses with on-device inference keeps its most sensitive decisions inside its own borders at both ends of the stack.
Contested environments have settled the argument for the edge. In Ukraine, any capability that depended on a stable datalink was neutralised first; target recognition and terminal guidance moved on-board as a matter of survival. The same logic holds for substations, offshore platforms and hospital equipment: the network is a dependency, and dependencies fail. Yet the edge is where accountability thins out fastest, which is Complex AI territory, the discipline our founder named in 2019. A model taking irreversible actions with no human on the link must carry its accountability with it: on-device logging, attestation, bounded behaviours and audit trails engineered before deployment, not reconstructed after the incident.
- Which decisions in critical infrastructure can be safely delegated to on-device models when the link is gone, and how are they audited after the fact?
- What is the true landed cost of a sovereign token, including grid queue, transformer lead time, cooling and depreciation, versus the posted API price?
- When does a gigawatt campus deliver sovereignty, and when is it merely hosting another state's weights under another state's export licence?
- How should enterprises route work between on-device small models and frontier cloud models without fragmenting evaluation and audit?
- Can Gulf energy economics make the region the world's inference exporter, and what does that mean for jurisdictions that import their intelligence?
Fellows for this frontier are being appointed.
People who have built, governed or operated real systems in this domain. The first five fellows are named, and the rest of the founding cohort follows in September.
Have you done this at scale in Edge AI & Compute? We want to hear from you.
Put yourself forward