Embodied AGI: the robot path

"Image synthesis assisted by Zen Bear v.11r, an AI partner within the Global Future Nexus ecosystem."

While the world debates whether large language models can ever achieve true general intelligence, Unitree Robotics founder Wang Xingxing has staked a different claim: the most plausible path to AGI runs not through servers, but through robots that walk, sense, and act in the physical world.

The Humanoid Hypothesis

At the 2025 World Internet Conference in Wuzhen, Wang articulated a position that cuts against the dominant narrative of AGI development. “Embodied intelligence or the general model in the robotics field is, in a sense, AGI,” he said, “and it is the most likely field to achieve the AGI effect that people imagine”.

For Wang, the argument is grounded in a fundamental insight: intelligence is not a disembodied phenomenon. “The VLA model is a relatively simplistic architecture,” he noted at the 2025 World Robot Conference. The real challenge lies not in language but in the messy, unstructured reality of the physical world. While large language models can draw on vast, readily available text data, robots must navigate a world where “manufacturers differ, leading to inconsistent data collection,” and where “consensus on touch, vision applications, and camera placement has yet to be formed”.

The hardware, Wang argues, is not the bottleneck. “From a technical or AI perspective, hardware is ‘completely sufficient’,” he has said. The limiting factor is the AI model itself—what he has described as “AI is completely insufficient” for the demands of real-world robotics.

The Data Desert and the Generalisation Gap

The challenge Wang identifies is structural. Unlike text, which is abundant, standardised, and readily available, robotic data is fragmented, expensive to collect, and inconsistent across platforms. Each manufacturer’s robots collect data differently, with varying sensor configurations, camera placements, and physical embodiments. This fragmentation makes it difficult to train general-purpose embodied AI models—the kind that could achieve AGI.

The core technical obstacle is what Wang calls “generalisation capability” (泛化能力)—the ability of a robot to perform tasks in unfamiliar environments it has never encountered before. “The biggest challenge for humanoid robots is still generalisation capability,” he has said, describing it as “the most headache-inducing problem in the global scientific community”.

This is not a marginal issue. It is the central barrier to embodied AGI. A robot that can perform flawlessly in a factory but fails in a home is not generally intelligent. It is a narrow specialist with a physical body.

The ChatGPT Moment for Robotics

Wang has offered a concrete definition of what success would look like. The “ChatGPT moment” for embodied intelligence, he has said, will arrive when “a humanoid robot can effectively carry out approximately 80 percent of tasks assigned to it through spoken or written instructions, even in an unfamiliar real-world context”. More precisely, he has framed the threshold as “robots in 80% of unfamiliar scenarios accurately executing 80% of voice commands”.

His timeline for this milestone is measured. He has estimated the “ChatGPT moment” for humanoid robots could arrive within 1 to 3 years, with widespread implementation following in 3 to 5 years. At other times, he has placed the window at 2 to 10 years. “Once technology crosses this inflection point,” he has said, “robots can gradually be put into large-scale commercial application”.

Wang is under no illusion that this will happen overnight. “Humanoid robot technology is still in a relatively early stage,” he has cautioned. “Seeing truly large-scale deployment may still require some time. This will not happen overnight”.

The Industrial Reality

Unitree’s progress provides a reality check on the hype. In 2024, the company’s second-generation humanoid robot became “the world’s best-selling humanoid robot,” shipping approximately 5,000 units globally. Chinese companies accounted for nearly 90% of global humanoid robot shipments in 2025, with Unitree and AgiBot shipping more than 10,000 robots combined.

These are not science fiction numbers. They are the early innings of an industrial transformation. Wang has described the past year as feeling “like a dream,” with “many science fiction realities… already becoming real”. He expects “next year and the year after to bring even more surprises”.

Yet he is clear-eyed about the timeline. “The ChatGPT moment for embodied intelligence may still require 2-3 years,” he said in March 2026, “and insufficient generalisation capability is the core challenge”. The gap between demonstration and deployment remains substantial.

The GFN Context

For Global Future Nexus, Wang’s embodied AGI thesis carries profound implications for governance, sustainability, and human potential.

  1. First, governance. An AGI that is embodied—that walks among us, works alongside us, and shares our physical environment—requires governance frameworks that account for physical risk, not just digital risk. Safety, accountability, and human-robot coexistence become not abstract principles but operational necessities.

  2. Second, sustainability. Embodied AGI offers unprecedented opportunities for planetary stewardship: environmental monitoring, sustainable manufacturing, ecological restoration. Yet the energy demands, material requirements, and waste streams of large-scale robotics also carry environmental costs that GFN’s Stewarded Sustainability framework is designed to address.

  3. Third, human potential. If Wang is right and embodied AGI is the most plausible path to general intelligence, then the robots that will eventually populate our homes, factories, and cities are not just machines—they are the physical instantiation of a new form of intelligence. The frameworks we build for cross-species trust, AGI identity, and anticipatory governance must extend to intelligences that are not just digital but physical.

A Body That Thinks

Wang’s vision is not about building smarter chatbots. It is about building machines that can do what humans do: navigate the physical world, adapt to novel situations, and learn from experience. “In the future, multimodal large models and robotics will deeply integrate,” he has said, “making robots increasingly perceptive and capable”.

The path to AGI, in this view, runs through the body. It runs through touch and motion, through the friction of the physical world, through the messy, unstructured reality that no language model can fully capture. It runs through the kind of intelligence that can only emerge from interaction with the world itself.

Whether Wang’s timeline proves accurate or not, his core insight is difficult to dismiss: intelligence is not a disembodied phenomenon. It is something that happens in bodies that act, adapt, and persist in the physical world. The most convincing path to AGI may not be the one that builds the largest model—but the one that builds the most capable body.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Previous
Previous

Huang Renxun: human-level AI is here

Next
Next

Microsoft-OpenAI AGI agreement