The scaling law vs. new paradigm debate
"Image synthesis assisted by Seedream 5.0 Pro, an AI partner within the Global Future Nexus ecosystem."
For four years, the AI industry lived by a single commandment: scale everything. More parameters, more data, more compute. But in 2026, that faith has shattered. As diminishing returns set in and the industry confronts a trillion-dollar reckoning, a fundamental debate has emerged: is scaling the path to AGI, or a dead end?
The Wall That Wasn't Supposed to Exist
The scaling law, first formalised by OpenAI in 2020, promised a simple formula: 10 times the parameters, 10 times the data, 10 times the compute equals 10 times better AI. It worked from GPT-3 to GPT-4. It doesn't work anymore.
Entering 2026, the scaling law hit a high wall: the returns from simply hoarding compute and piling on data have begun to diminish. Spending tens of billions of dollars to train the next generation of super-models yields performance improvements that are far from revolutionary. The industry is feeling it: larger parameters and more chips are no longer delivering proportional intelligence gains.
Ilya Sutskever, OpenAI's former chief scientist and one of the architects of the scaling era, has admitted that traditional unsupervised pretraining has reached its limit. His exact words: these models "generalise dramatically worse than people". Yann LeCun, true to form, has argued that no amount of scaling current large language models will reach true AGI.
The Reckoning: A Trillion-Dollar Question
The financial stakes are staggering. Meta, Amazon, Microsoft, Google, and Tesla have spent $560 billion on AI capital expenditure since early 2024—and generated just $35 billion in AI revenue. That is a 16-to-1 spend-to-revenue ratio. AI-related spending now accounts for 50% of US GDP growth, and the White House has admitted that a reversal would risk recession.
Meanwhile, a bombshell revelation has further shaken the industry. In July 2026, a DeepMind researcher revealed that OpenAI's original scaling law paper contained a fundamental bug that likely led the industry to waste vast amounts of compute on oversized, undertrained models. As one commentator put it: "A bug, burning two years". The scaling law was never Newton's laws—it was always an empirical curve.
The New Paradigms: What Comes After Scaling?
As the industry moves into what WAIC 2026 called the "post-scaling era", several alternative paradigms are emerging:
World Models and Physical Intelligence are perhaps the most prominent alternative. Instead of predicting the next token, world models predict the "next state"—how a system evolves in physical reality. This approach addresses the fundamental critique that language is not the origin of intelligence, but a late-stage compression layer built on top of far more primitive capabilities. Yann LeCun left Meta to raise $1.03 billion to train world models. Google's Gemini Omni and other major players are betting on this direction. As one industry observer noted, there is a "new consensus among AI giants: relying solely on text seems insufficient for achieving AGI".
Structure Learning offers another path. Instead of merely accumulating parameters, structure learning focuses on enabling artificial agents to learn a good model of the world—including a good model of human preferences. A 2026 paper argues that active causal structure learning is a necessary component for building AGI agents. This approach aims to build systems that understand causality, not just correlation.
Neuro-Symbolic AI integrates neural networks with symbolic reasoning systems. Neural networks provide pattern recognition; symbolic systems provide reasoning, verification, and compositionality. This hybrid approach addresses the reasoning gap that pure LLMs cannot bridge. It has been identified as a 600% growth area and is emerging as a critical path forward for systems that need to reason independently and enforce constraints reliably.
Test-Time Compute represents a different kind of scaling—scaling at inference rather than training. OpenAI's o1/o3 models allocate additional computation during reasoning. Stanford's s1 model, trained on just 1,000 examples, beat o1-preview by 27% on competition math using budget forcing. This approach proves that intelligent compute allocation can beat brute scale.
Small Language Models are the efficiency play. Models with 1 to 10 billion parameters are matching GPT-4 in narrow domains, running on-device, preserving privacy, and costing 10 times less. They don't require hyperscale infrastructure, making AI more accessible and sustainable.
Modular and Multi-Agent Architectures are redefining how we think about intelligence. Instead of a single monolithic model, modular systems combine specialists. Research has demonstrated that merging specialists offers a complementary path toward AGI-level systems. The WAIC 2026 consensus is that the industry is moving from "crude model competition" to an "agentic productivity era" centred on deployment, not just capability.
The GFN Context: Governing the Transition
For Global Future Nexus, the scaling debate is not merely technical—it is a governance imperative. The shift from "scale everything" to "what actually works" creates a moment of institutional flux, when the old certainties have collapsed and the new ones have not yet taken shape. This is precisely the kind of moment when anticipatory governance matters most.
The debate also has profound implications for planetary sustainability. The $560 billion spent on AI infrastructure, the energy consumption of trillion-parameter models, and the waste heat generated by compute-intensive training all have environmental costs that GFN's sustainability framework is designed to address. If new paradigms—world models, structure learning, small models—can deliver intelligence with less compute, they may also deliver intelligence with less environmental damage.
The most significant insight may be that AGI will not emerge from a single paradigm. It will emerge from the convergence of multiple approaches: world models that understand physics, neuro-symbolic systems that reason, modular architectures that coordinate, and learning systems that build causal understanding. The governance frameworks we build must be architecture-agnostic—capable of adapting to a landscape where the only certainty is change.
The scaling era is ending. The research era is beginning. And the question is no longer whether we will scale—but what we will build instead.
Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)
Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)