AGI and the physical world understanding
"Image synthesis assisted by Microsoft Copilot, an AI partner within the Global Future Nexus ecosystem."
From next-token prediction to physical reasoning, the path to true intelligence requires something current AI systems fundamentally lack: a genuine understanding of how the physical world works.
The current trajectory of artificial intelligence, defined by the brute-force scaling of Transformer-based architectures, has brought remarkable capabilities but also fundamental limitations. The core problem is what researchers call the Embodiment Gap: the disconnect between digital prediction (next-token generation) and physical understanding, resulting in hallucinations and a lack of spatial intelligence. As one analysis from the 2026 Zhiyuan Conference notes, current embodied AI systems suffer from a "demonstration-level to practical-level" gap—they can demonstrate tasks in controlled environments but fail when faced with real-world complexity.
This gap is not merely a technical inconvenience. Without a grasp of thermodynamics, electromagnetism, optics, and causal relationships, AI systems cannot reliably interact with the physical world, plan long-horizon tasks, or avoid catastrophic errors in safety-critical applications.
Why LLMs Alone Are Not Enough
Large Language Models are statistical pattern-matchers. They predict the next token based on patterns in their training data. They do not simulate physical reality. As research from the Beijing Institute for General Artificial Intelligence observes, "A three-year-old child can instinctively understand the meaning of a wave, can foresee a cup sliding off the edge of a table—while machines struggle to translate these visual signals into a deep understanding of physical laws".
This limitation has been empirically documented. The Physics-RW benchmark, constructed from real-world videos across mechanics, thermodynamics, electromagnetism, and optics, revealed that current general world models exhibit "limited physical reasoning capabilities". A separate benchmark, NovPhy, found that while humans can detect and adapt to novel physical scenarios, "some agents, even with good normal task performance, perform significantly worse when there is a novelty".
The fundamental issue is ontological. As a recent critique of world models argues, LLMs simulate language patterns, not physics. A world model must simulate all "actionable possibilities of the real world for purposeful reasoning and acting"—a task that requires hierarchical, multi-level representations grounded in sensory experience, not just text .
World Models as the Necessary Path
The emerging consensus is that World Models—systems that build internal representations and predictions of the external world—are the bridge to genuine physical understanding . As Fei-Fei Li has argued, "understanding 3D worlds, generating 3D worlds, reasoning in 3D worlds, and acting in 3D worlds are fundamental problems of artificial intelligence." Without spatial intelligence, AGI remains incomplete.
The PAN architecture (Physical, Agentic, and Nested) proposes a general-purpose world model based on hierarchical, multi-level, and mixed continuous/discrete representations, enabling agents to perform "thought experiments" by simulating actions and plans in complex scenarios—including counterfactual ones.
For Global Future Nexus, the shift from language to physics is central to the mission of responsible AGI integration. AGI that cannot understand the physical world cannot steward planetary sustainability—it cannot model climate systems, optimise resource flows, or interact safely with physical infrastructure. GFN's work on embodied intelligence, cross-species trust, and sustainable AGI deployment depends on systems that are grounded in reality, not just fluent in text.
As one researcher put it: "the real world is not a language game". The future of AGI depends on whether we can teach machines to understand physics, not just predict tokens.
Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)
Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)