The ARC-AGI benchmark milestone
"Image synthesis assisted by Seedream 5.0 Pro, an AI partner within the Global Future Nexus ecosystem."
In early 2026, the AI community quietly passed a threshold that would have seemed impossible just a few years earlier. On the ARC-AGI-2 benchmark—designed specifically to resist memorisation and test fluid intelligence—frontier AI systems surpassed average human performance for the first time. Yet even as the milestone was celebrated, Yann LeCun was already cautioning that this achievement, while impressive, does not mean the LLM path leads to AGI.
The Benchmark That Refuses to Be Gamed
The ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) benchmark was created by François Chollet—creator of Keras and a leading AI researcher—with a deceptively simple premise: to test whether AI can reason like a human. Unlike traditional benchmarks that can be "solved" through memorisation or statistical pattern matching, ARC-AGI presents novel grid-transformation puzzles where the system must infer the underlying rule from just a few examples. There is no massive training set; every task is novel. The benchmark is a test of fluid intelligence—the ability to adapt to new situations, not just recall patterns.
For years, this benchmark remained stubbornly out of reach. GPT-4.5 scored 0.0% on ARC-AGI-2. The winning open-source entry in the ARC Prize 2025 reached just 24.03%. In May 2025, the strongest baselines achieved approximately 3%. The benchmark was designed to be "easy for humans, hard for AI", and for a long time, it lived up to that billing.
The Threshold Is Crossed
January 2026 marked a turning point. OpenAI's Greg Brockman announced that GPT-5.2 had crossed the human baseline on ARC-AGI-2. The Thinking version scored 52.9%, while the Pro version reached 54.2%—a nearly threefold improvement over the previous generation's 17.6%.
The most dramatic result came from Poetiq (GPT-5.2X-High), a meta-system that automatically orchestrates frontier models through software architecture rather than training larger models. Poetiq achieved 75% accuracy on ARC-AGI-2, surpassing the human average of 60% by 15 percentage points. Crucially, this was achieved without any additional training or fine-tuning of the base model. As Poetiq demonstrated, intelligent software architecture can extract substantially more reasoning capability from existing models—a finding with profound implications for the future of AI development.
By February 2026, Google's Gemini 3 Deep Think had pushed the frontier even further, achieving 84.6% on ARC-AGI-2—a record-breaking score that placed it well into "superhuman" territory. The model also scored 96.0% on ARC-AGI-1, while Claude Opus 4.6 reached 68.8% on ARC-AGI-2.
The Costs of Progress
What makes these achievements remarkable is not just the performance but the economics. The cost of achieving high performance on ARC-AGI-2 plummeted 390-fold in a single year—from o3's $4,500 per task to GPT-5.2's $12 per task. Open-weight models like Thinking Machines' Inkling now score 36.5% on ARC-AGI-2 at an average cost of just $0.64 per task.
Yet the milestone also reveals a deeper challenge. The performance gap between humans and AI across ARC-AGI versions is stark:
ARC-AGI-1: AI systems reach 93.0% (Claude Opus 4.6)
ARC-AGI-2: Performance falls to 68.8%
ARC-AGI-3: Performance plummets to 13%
Humans maintain near-perfect accuracy across all versions. As one analysis concluded, "performance degradation across versions is consistent across all paradigms: program synthesis, neuro-symbolic, and neural approaches all exhibit 2-3x drops from ARC-AGI-1 to ARC-AGI-2, indicating fundamental limitations in compositional generalization".
LeCun's Caution: The Limits of the Milestone
The ARC-AGI-2 milestone, while significant, does not settle the deeper debate about the path to AGI. Yann LeCun has spent considerable energy arguing that pure language modelling is a dead end and that something model-based—agents that build and reason over internal simulators of their environments—is what serious progress toward AGI actually requires.
LeCun has compared the LLM path to "climbing a tall tree to reach the moon". He argues that even as models achieve impressive scores on benchmarks, they lack the foundational capabilities required for genuine understanding: causal reasoning, physical intuition, and the kind of common sense that a four-year-old child possesses. He has called the idea that "the path to superintelligence just requires training LLMs with more synthetic data, hiring thousands of people to post-train your system, and inventing new reinforcement learning tricks" "complete nonsense".
The release of ARC-AGI-3 in 2026 reinforces LeCun's critique. Unlike its static predecessor, ARC-AGI-3 requires agents to explore unknown environments, acquire their own goals, and plan long-horizon actions without instructions. On this benchmark, even the most advanced systems score in the low single digits: GPT-5.6 Terra averages just 0.8%. To solve ARC-AGI-3, AI must demonstrate exactly the capabilities—exploration, autonomous goal-setting, adaptive planning—that LeCun has argued are missing from LLM-only approaches.
The GFN Context: Milestones and Governance
For Global Future Nexus, the ARC-AGI-2 milestone carries a dual significance. On one hand, it demonstrates that AI systems can match or exceed human performance on tasks designed to test general reasoning—a step toward the kind of broad capability that AGI requires. On the other hand, the precipitous drop on ARC-AGI-3 underscores that fundamental gaps remain.
The ARC-AGI milestone also highlights a governance challenge: if AI systems can outperform humans on reasoning benchmarks while still lacking the foundational capabilities that LeCun identifies, how do we assess readiness for real-world deployment? The benchmark may measure one dimension of progress, but it cannot measure everything that matters. GFN's frameworks for AGI identity, cross-species trust, and anticipatory governance must account for the gap between benchmark performance and genuine understanding—a gap that the ARC-AGI series so vividly reveals.
Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)
Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)