OpenAI's GPT-5 family: a step toward AGI?
"Image synthesis assisted by Qwen Image 3.0, an AI partner within the Global Future Nexus ecosystem."
From “PhD-level expert” to 74% outperformance of human professionals, OpenAI's GPT-5 family represents the most significant step toward artificial general intelligence yet—while its creator cautions that the journey is far from complete.
A New Kind of Intelligence
In late 2025, OpenAI released GPT-5, a model that its creators describe not as a language model but as a “General Reasoning Engine.” As Sam Altman put it at launch: “This is the first time that it really feels like talking to a PhD-level expert in any topic.” The distinction is fundamental: where GPT-4 and its predecessors operated through probabilistic text generation, GPT-5 was engineered to perform logical task solving through genuine reasoning.
The architecture reflects this shift. GPT-5 integrates System 2 thinking—slow, deliberate, logical reasoning—alongside the intuitive System 1 processing of earlier models. Before outputting an answer, the model executes thousands of steps of internal trial, error, verification, and self-correction through internal chain-of-thought processing. It explores multiple solution paths in parallel, evaluates their feasibility, and selects the optimal path. The result is a system that, in mathematics Olympiads, top-tier programming challenges, and legal case analysis, now outperforms 99% of human experts.
GPT-5 is also natively multimodal—all modalities interact directly within the same neural network space, rather than converting images or audio to text before processing. It understands hours of video content, captures subtle facial expressions and causal relationships, reduces voice dialogue latency to under 150 milliseconds, and can generate runnable front-end code from hand-drawn sketches. The Pro version supports a 10 million-token context window—enough to read an entire library of legal texts or years of scientific papers in a single pass.
The Capability Frontier
By July 2026, the GPT-5 family had evolved into a three-tiered system: Luna (budget, $1 per million input tokens), Terra (mid-range, $2.50), and Sol (flagship, $5). All three share a February 2026 knowledge cutoff, a million-token context window, and 128,000 maximum output tokens.
The performance metrics are unprecedented. In internal benchmark GDP-Val, spanning over 40 knowledge domains, GPT-5.2 matches or outperforms human experts in 74.1% of tasks. As Altman stated: “AI is a colleague you give an hour of work to—and 3 out of 4 times it's better than a human. Three years ago: ~0%. Today: 74%.” GPT-5.2 Thinking achieves 70.9% on GDPVal, compared to GPT-5's 38.8%. It scores 92.4% on GPQA Diamond scientific questions and 99.4% on HMMT mathematics competition problems. On ARC-AGI-2, a benchmark designed to resist memorisation and test fluid intelligence, GPT-5.2 reaches 52.9%.
The latest iteration, GPT-5.6 Sol, sets new records in long-running agentic performance. On Agents' Last Exam, an evaluation of professional workflows across 55 fields, Sol achieves 53.6—outperforming Anthropic's Claude Fable 5 by 13.1 points. With multi-agent orchestration, Sol can coordinate four parallel agents to tackle complex tasks simultaneously. It introduces “max reasoning effort” and “ultra” modes, allowing the model to spend more time on reasoning, explore alternatives, and self-correct—or default to coordinating multiple agents for complex tasks.
A Giant Step—Still Short of AGI
Yet Altman is careful to distinguish capability from consciousness. “This is not a model that continuously learns as it is deployed from new things it finds, which is something that, to me, feels like it should be part of an AGI,” he cautioned. “But the level of capability here is a huge improvement.”
The distinction matters. AGI, in Altman's framing, requires continuous learning—the ability to adapt and grow from new experiences without retraining. GPT-5 does not do this. It is a snapshot of intelligence, not a living intelligence. “There is still work to be done to achieve the kind of artificial general intelligence that thinks the way people do.”
Altman has also acknowledged that “AGI kind of already happened. We're arguing about definitions.” This is not a contradiction. It is a recognition that the goalposts have shifted. In three years, the question has moved from “Can AI pass the Turing test?” to “Can AI run countries better than humans?” The new target is Superintelligence—AI that could outperform the best human CEOs or presidents.
Altman expects “orders of magnitude more gains” to come on the path toward AGI. “Obviously... you have to invest in compute at an eye-watering rate to get that, but we intend to keep doing it.”
The GFN Context
For Global Future Nexus, the GPT-5 family represents both the promise and the governance challenge of the AGI era. A model that outperforms human experts in 74% of tasks yet cannot learn continuously from experience is a reminder that capability and understanding are not the same thing. The gap between reasoning and consciousness, between performance and wisdom, is precisely the gap that GFN's frameworks for AGI identity, cross-species trust, and anticipatory governance are designed to address.
GPT-5 is not AGI. But it is the clearest signal yet that the journey from narrow AI to general intelligence is no longer theoretical—it is engineering. The question is no longer whether we will build systems that reason like experts. The question is whether we will build the governance frameworks to ensure that when they do, they serve human flourishing rather than human obsolescence. The window for that governance is closing as fast as the models are advancing.
Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)
Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)