AGI's path: agents, architecture, training

"Image synthesis assisted by Qwen Image 3.0, an AI partner within the Global Future Nexus ecosystem."

From 35-billion-parameter models that rival trillion-scale systems to self-evolving agents that improve themselves through recursive loops, 2026 has witnessed a fundamental reorientation in the quest for artificial general intelligence—a shift from scaling parameters to scaling intelligence itself.

The Year the Agent Took Centre Stage

In January 2026, Sequoia Capital declared that 2026 would be the "Year of Long-Horizon Agents," arguing that AI capable of operating autonomously over extended timeframes "is functionally equivalent to AGI". This proclamation captured a broader industry shift: the conversation had moved from parameter counts to practical capability.

The evidence was immediate. In July 2026, Shanghai AI Lab released Agents-A1, a 35-billion-parameter Mixture-of-Experts model that reaches trillion-parameter-level performance—not by scaling parameters, but by scaling the agent horizon. The model handles trajectories averaging 45,000 tokens across six heterogeneous domains—long-horizon search, scientific research, instruction following, and engineering tasks—and achieves state-of-the-art results on benchmarks including GAIA (96.0%) and BrowseComp (75.5%). As the research team put it: the goal is to teach models "more persistent, verified work habits, rather than merely expanding their parameter scale".

Agents-A1's architecture represents a paradigm shift. It uses a three-stage training pipeline: full-domain supervised fine-tuning to establish general agentic capabilities, domain-level teacher models capturing specialised expertise in search, science, instruction following, and tool use, and multi-teacher on-policy distillation that unifies these capabilities into a single deployable model. The result is a system that can decompose complex tasks, plan ahead, and adapt strategies based on intermediate results—capabilities that previously required models an order of magnitude larger.

The Architecture of Intelligence

Alongside agentic breakthroughs, 2026 has seen fundamental architectural innovation. Google DeepMind's Co-Scientist, published in Nature in May 2026, introduced a multi-agent system built on Gemini that iteratively generates, debates, and evolves novel hypotheses for complex scientific problems. Its architecture comprises specialised agents across three phases: Generation agents propose hypotheses; Reflection and Ranking agents conduct "idea tournaments" through simulated scientific debates; Evolution and Meta-review agents refine and synthesise the most promising directions. The system has already identified novel drug-repurposing candidates, including one that blocked 91% of a scarring-linked response in lab tests.

On the hardware-software frontier, the AgentOS framework, published in February 2026, reimagines the LLM as a "Reasoning Kernel" governed by structured operating system logic, mapping classical OS abstractions—memory paging, interrupt handling, process scheduling—onto LLM-native constructs. This reframing asserts that "the next frontier of AGI development lies in the architectural efficiency of system-level coordination".

World models have also emerged as a critical architectural pillar. A comprehensive review from Shanghai AI Lab identified world models as "the key path to physical AGI," proposing a trinity architecture: Agent executing tasks, Evaluator assessing trajectories, and World Model absorbing interaction data and generating new tasks. The review defines a world model as "a compressed modelling of the physical world's state transition process under finite computational resource constraints", encompassing three functional components: a Renderer (observations to perceptible results), a Simulator (predicting how states evolve), and a Planner (evaluating counterfactual futures). Major players have taken notice: Kunlun Tech declared 2026 "the first year of world models", releasing Riemann-1.0 at WAIC 2026.

The Training Revolution

Training methodologies have undergone equally dramatic evolution. Tsinghua University's SEED (Self-Evolving On-Policy Distillation) framework addresses a fundamental limitation of reinforcement learning: outcome-based RL can only judge task success or failure, failing to identify which intermediate actions should be reinforced. SEED solves this by having the model summarise completed trajectories into "hindsight skills"—successful trajectories yield reusable strategies; failed ones yield correction rules—then distils these skills back into the policy. Results are striking: SEED improved performance on ALFWorld by 45.9 percentage points over GRPO.

The Red Queen Gödel Machine, developed by Cambridge researchers with collaborators from NVIDIA and Flower Labs, tackles the evaluation ceiling that plagues self-improving AI. When a self-improving agent learns everything a fixed evaluator can distinguish, improvement stalls. The solution: let both agent and evaluator evolve together. "As the agent gets better, the evaluation also gets harder, and the bar keeps rising". The system searches through many possible versions of an agent while simultaneously improving the evaluator that judges them, creating a self-improving loop.

Recursive self-improvement has also moved from theory to practice. Tencent's Hyra-1.0, released in July 2026, is a research agent capable of recursive self-improvement, functioning like a research scientist: proposing hypotheses, conducting experiments, synthesising lessons, and iterating. Former Meta FAIR research director Tian Yuandong's startup Recursive has demonstrated automated AI research systems achieving state-of-the-art results on NVIDIA's GPU kernel optimisation leaderboard.

The GFN Context

For Global Future Nexus, these breakthroughs in agents, architecture, and training define the trajectory of AGI development that governance frameworks must address. The emergence of long-horizon agents capable of autonomous multi-step reasoning, world models that bridge digital and physical intelligence, and training methods that enable recursive self-improvement all point to a future where AGI is not a single event but a continuous process of capability expansion. GFN's work on AGI identity, cross-species trust, and anticipatory governance must evolve in parallel with these technical advances—ensuring that as agents become more capable, they remain aligned with human values and planetary sustainability.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Previous
Previous

Safe AGI: the center for AI safety

Next
Next

The definition of AGI: a new consensus