The cognitive versatility question

"Image synthesis assisted by Mai Image 2.5, an AI partner within the Global Future Nexus ecosystem."

From a "jagged intelligence" profile that excels at Olympiad math but fails basic common-sense tasks, to a 57% AGI score that masks catastrophic deficits in long-term memory, the question of whether AI can match human cognitive abilities reveals a profound disconnect between surface-level performance and genuine understanding.

The Measurement Challenge

For decades, Artificial General Intelligence has remained shrouded in fuzzy definitions — "human-like intelligence," "general-purpose reasoning" — that have made it nearly impossible to track meaningful progress. In October 2025, a coalition of AI leaders including Yoshua Bengio, Dan Hendrycks, and Eric Schmidt published a landmark paper that cut through the ambiguity: AGI is "an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult".

This definition rests on two critical dimensions: versatility (the breadth of skills across multiple domains) and proficiency (the depth of competence in each). To operationalise this, the researchers turned to the Cattell-Horn-Carroll (CHC) theory — the most empirically validated model of human intelligence — and decomposed general intelligence into ten core cognitive domains, including reasoning, memory, perception, and processing speed.

The result: an AGI score that offers a concrete, quantifiable target. GPT-4 scored 27%; GPT-5 reached 57%. Yet this number conceals more than it reveals.

The Jagged Profile

What emerges from the CHC framework is not a uniform intelligence but a "jagged" cognitive profile. In knowledge-intensive domains — general knowledge, mathematics, reading and writing — GPT-5 performs at or near human levels. But the picture changes dramatically in areas that form the foundations of genuine cognition.

The domain scores are stark:

  • General Knowledge: 9% (GPT-5)

  • Reading and Writing: 10%

  • Mathematics: 10%

  • On-the-Spot Reasoning: 7%

  • Working Memory: 4%

  • Long-Term Memory Storage: 0%

  • Long-Term Memory Retrieval: 4%

  • Visual Processing: 4%

  • Auditory Processing: 6%

  • Processing Speed: 3%

The zero in Long-Term Memory Storage is not a minor gap — it is a catastrophic failure in one of the most foundational cognitive faculties. As the authors note, current AI systems have "critical deficits in foundational cognitive machinery" that are masked by their proficiency in knowledge-intensive domains.

The Coherence Challenge

Some researchers argue that even the 57% score may overstate progress. A coherence-based critique points out that the arithmetic mean assumes compensability — that exceptional performance in some areas can offset severe deficiencies in others . Yet genuine general intelligence requires "coherent sufficiency": balanced competence across all essential faculties.

The argument is grounded in both psychometric evidence and systems theory. In human cognition, abilities are interdependent — reasoning depends on working memory, perception constrains abstraction, and learning relies on durable long-term memory . In complex engineered systems, overall capability is limited by the weakest component.

When the arithmetic mean is replaced with a coherence-aware measure that penalises imbalance, the coherence-adjusted score for GPT-5 drops to 24% — revealing that both systems "remain far from general competence despite high arithmetic scores".

The Structural Incompatibility

A separate investigation of nine frontier models using the CHC framework revealed a deeper problem: applying human psychometric frameworks to transformer architectures may constitute a "fundamental category error". Models achieving above-average human IQ scores simultaneously exhibited binary accuracy rates approaching zero on crystallised knowledge tasks, with an overall correlation of just 0.175 between conceptual accuracy and exact-match scoring.

The "Verbosity Paradox" compounds the issue: while human psychometric tests expect minimal, precise responses, large language models are optimised through RLHF to provide comprehensive, explanatory answers. The result is a systematic measurement failure wherein models providing conceptually perfect answers receive zero scores due to response format misalignment.

The GFN Imperative

For Global Future Nexus, the cognitive versatility question is not an academic exercise — it is a governance imperative. If AGI systems possess a "jagged" intelligence profile — capable of PhD-level reasoning in some domains while failing at simple common-sense tasks that a child could perform — then the frameworks for trust, accountability, and coexistence cannot assume uniform capability.

The zero in Long-Term Memory Storage means that these systems cannot learn from experience in any meaningful sense. The coherence challenge means that aggregate scores mask catastrophic failures. The structural incompatibility means that our tools for measuring intelligence may be fundamentally ill-suited for the substrate we are evaluating.

As one analysis concluded, "the real question is not whether LLMs can match human performance on specific tasks, but whether they possess the cognitive machinery that makes human intelligence coherent, adaptable, and reliable" . That machinery — long-term memory, working memory, on-the-spot reasoning — remains conspicuously absent. The 57% score is not a milestone. It is a diagnostic of how far we still have to go.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Next
Next

AGI and the future of innovation