AGI's computational infrastructure

"Image synthesis assisted by Seedream 5.0 Pro, an AI partner within the Global Future Nexus ecosystem."

The race to AGI is a race to build the world's most powerful computing infrastructure. Beneath every frontier model, every reasoning breakthrough, and every agentic workflow lies a rapidly evolving stack of silicon, systems, and scale—a new kind of computing architecture that treats intelligence not as a product of individual processors, but as an emergent property of integrated systems.

The Silicon Frontier

The foundation of AGI infrastructure is purpose-built silicon. In 2026, the chip market is no longer a single-player game. AMD launched its Helios rack-scale system, connecting 72 Instinct MI455X GPUs with 31 TB of unified HBM4 memory and delivering 2.9 exaflops of FP4 compute—all on an open-standards Ethernet network that challenges Nvidia's proprietary interconnects. Anthropic has committed to up to 2 gigawatts of MI450-series GPUs, while OpenAI has been running GPT-class workloads on Helios for months.

Arm entered the data centre CPU market with its AGI CPU, targeting the exploding demand for CPU cores in agentic AI workloads. Where GPUs generate tokens, CPUs orchestrate them: managing agent sandboxes, coordinating tool calls, and executing the code that agents generate. In a single rack, Arm's liquid-cooled configuration packs over 45,000 cores—a scale that reflects the shift from single-model inference to persistent, multi-agent workflows.

The shift from training to inference has redefined the role of the CPU. In the agentic era, the CPU does not merely support the GPU—it "operationalises intelligence," handling the query routing, tool use, and orchestration that transforms raw tokens into meaningful action. The CPU is no longer a support processor; it is the control plane of intelligence.

The Hyper-Scale Clusters

Above the chip lies the cluster. Google's seventh-generation Ironwood TPU delivers a 10x peak performance jump over its predecessor, with 192GB of HBM3e memory and 7.4 TB/s bandwidth. In 9,216-chip "superpod" configurations, Ironwood enables the kind of massive key-value caches required for million-token reasoning models. Anthropic is deploying a million Ironwood TPUs—a multi-hundred-billion-dollar investment that will push compute clusters into the 1+ gigawatt power range.

China has answered with the Sugon 8000—the first fully domestic 100,000-accelerator AI supercluster. Built on a "super-intelligence fusion" architecture that unifies high-precision scientific computing with low-precision AI training, the cluster connects over 1600 km of cabling and 30 billion electronic components into a single coherent system. Its 100,000-accelerator deployment marks the arrival of China's sovereign AI infrastructure at the gigawatt scale.

The PC AI Superchip

For the first time, AGI infrastructure is arriving on the desktop. Nvidia's RTX Spark combines a Blackwell RTX GPU and a 20-core Grace CPU, unified through NVLink-C2C into a single superchip delivering 1 petaflop of AI performance. With 128 GB of unified memory, the chip can run 120-billion-parameter models locally, with million-token contexts, entirely on-device. The OpenShell runtime lets users define what their AI agents can do and keeps sensitive data on-device. The era of "personal AGI" is beginning—not at a data center, but in a laptop.

The System Problem

The most fundamental shift is architectural. AI infrastructure is no longer about individual processors—it is about rack-scale systems where compute, memory, interconnect, networking, storage, and software must be engineered together. Nvidia's Vera Rubin platform spans seven co-designed chips and five rack-scale configurations, treating the entire rack as a single supercomputer. The Vera CPU rack alone supports over 22,500 concurrent agent sandbox environments—a scale that reflects the persistence and parallelism of agentic AI .

Yet the system problem has a dark side. A KAIST study found that AI agents—which repeatedly call models, execute tools, and coordinate across multiple components—can consume up to 136 times more energy per request than standard generative AI. GPU utilisation efficiency can drop by over 50% as components wait on each other. The system-level inefficiencies are as consequential as the silicon-level efficiencies.

The GFN Context

For Global Future Nexus, the AGI infrastructure stack is both an opportunity and a governance challenge. The immense energy demands of future AGI must be met with sustainable solutions. The concentration of compute power raises questions of access and equity. The shift from cloud to edge—from billion-dollar clusters to personal superchips—creates new possibilities for decentralised intelligence, but also new vectors of risk.

The architectural truth is clear: AGI is not built in a single chip. It is built in a stack—and the stack is being assembled at a scale that will define the 21st century. The question is whether we will build it wisely, sustainably, and in service of human flourishing.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Previous
Previous

The moving goalpost: why we keep redefining AGI

Next
Next

Musk's AGI timeline acceleration