The "PhD-level" promise and its limits
"Image synthesis assisted by Qwen Image 3.0, an AI partner within the Global Future Nexus ecosystem."
When OpenAI CEO Sam Altman declared that GPT-5 offered "PhD-level expertise" in any domain, he made a claim that resonated far beyond Silicon Valley. But Demis Hassabis, the Nobel laureate leading Google DeepMind, delivered a sharp rebuttal: calling today's AI "PhD-level" intelligence is, in his words, "nonsense." While AI may exhibit flashes of expertise, it lacks the consistency, creativity, and genuine understanding that defines a human doctorate.
The Claim and the Counterclaim
OpenAI's definition of "PhD-level" AI originates from the company's five-level AGI roadmap. Level 2, the "Reasoner," is described as a system "capable of performing basic problem-solving tasks at a level comparable to a human with a PhD education, albeit without access to any tools". OpenAI's CTO had previously suggested that such capabilities could emerge within 18 months.
Altman's framing in 2025 positioned GPT-5 as having achieved this threshold. At the launch event, he declared: "GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD-level expert".
Hassabis's critique is structural. In a September 2025 interview at the All-In Summit, he argued that current AI systems may possess "certain PhD-level capabilities in particular sub-tasks, but not comprehensive or holistic capabilities." True PhD-level intelligence, Hassabis maintained, means performing consistently at the highest level "across all dimensions"—a standard current systems fail to meet.
What "PhD-Level" Truly Requires
The gap between AI's performance and genuine doctoral-level intelligence is illuminated by three key dimensions.
Breadth vs. Depth. A human PhD represents deep expertise within a field, combined with broad competence across adjacent domains. Frontier models may excel at narrow tasks, but they lack the integrated, multi-faceted reasoning required for genuine professional work.
Consistency. As Hassabis observes, a true "PhD-level" system should not fail on elementary tasks. Yet current AI systems routinely make "stupid mistakes on basic high-school math or simple counting just by altering the phrasing of the question". When an AI that can solve Olympiad-level mathematics fails to correctly count letters in a word, the gap between capability and reliability becomes starkly visible.
Creativity and Problem Formulation. Perhaps the most profound gap is the inability to pose novel questions. Hassabis identifies this as a critical threshold: "Under a given proposition, today's AI can go prove it or solve it, but it cannot come up with entirely new conjectures, hypotheses, or theories". The ability to formulate a research question is the defining characteristic of doctoral work, and it remains beyond AI's reach.
Beyond Knowledge to Understanding
The debate exposes a deeper truth: passing doctoral-level exams is not the same as possessing doctoral-level intelligence. Hassabis has proposed a concrete test: limit an AI's knowledge to 1901 and see if it can derive special relativity by 1905—a genuine creative leap, not a recombination of existing ideas. By this standard, today's AI fails entirely.
Gartner analysts have similarly noted: "Despite its strengths, GPT-5 doesn't deliver artificial general intelligence. It doesn't autonomously learn, lacks architectural vision, and still requires human oversight for mission-critical decisions". The system is a refinement, not a transformation.
The Governance Implications
For Global Future Nexus, the PhD-level debate is not an academic squabble. It is a governance question. If AI is marketed as PhD-level but functions as a highly capable but unreliable assistant, the risks of over-deployment are severe: unverified scientific claims, misdiagnoses in medicine, and a creeping erosion of professional standards.
The debate also highlights a broader governance gap: definitions of AGI capability are being set by companies competing for investment and market share, not by independent bodies. Without a shared, testable framework, claims of PhD-level intelligence remain claims, not facts. Hassabis has estimated that achieving PhD-level comprehensive capabilities may require five to ten more years, and perhaps one or two key breakthroughs that have not yet arrived.
The PhD-level promise is real—but it is also limited. The gap between AI's performance and genuine understanding is not a failure of the technology. It is a diagnostic, telling us where the science must go and where governance must focus.
Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)
Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)