The numbers behind the intelligence: understanding AI's scale

"Image synthesis assisted by Qwen Image 3.0, an AI partner within the Global Future Nexus ecosystem."

When GPT-3 arrived in 2020 with 175 billion parameters, the number was the headline. Bigger meant smarter. Five years later, that metric is effectively dead, yet a far richer vocabulary of figures now defines the AI landscape. Understanding these numbers is essential for grasping where we stand, how fast we are moving, and what the trajectory toward AGI and superintelligence actually looks like.

Parameters: The Measure That Changed

A model's parameters are the numerical settings it learns from data to recognize patterns and generate responses. In the dense-model era, total parameters directly indicated size and capability. Today, architecture matters more than headline numbers.

The shift began with Mixture-of-Experts (MoE). DeepSeek V3 has 671 billion total parameters but activates only 37 billion per token. A model with 1.6 trillion total parameters may use just 49 billion per token. The "billions of parameters" headline quietly stopped meaning anything around 2024. Dense models fire every parameter on every token — a 405 billion parameter model uses all 405 billion. MoE models split parameters into specialized experts, routing each token to a small fraction of the total.

Notable Models and Their Architectures

The range of current frontier models illustrates how much architecture matters:

  • DeepSeek V3 uses a Mixture-of-Experts architecture with 671 billion total parameters, activating only 37 billion per token. It achieved 93.1% on the AIME 2025 benchmark.

  • Qwen3-235B-A22B also uses MoE with 235 billion total parameters and 22 billion activated per token, scoring 81.5% on AIME 2025.

  • Grok-3 is estimated at approximately 2.7 trillion parameters with undisclosed active per token, achieving 93.3% on AIME 2025.

  • GPT-4 was estimated around 1.8 trillion total parameters with approximately 280 billion active per token.

  • Claude Opus 4.6 is estimated at approximately 5 trillion total parameters, though active per token remains undisclosed.

The contest has flipped: "biggest model possible" became "smallest model at this quality." GPT-4 was estimated around 1.8 trillion parameters; GPT-4o landed near 200 billion with comparable performance .

Compute: The New Currency

Effective compute grows by approximately 10× per year — a compound effect of hardware improvements, investment scale, and algorithmic efficiency. DeepMind's 2026 report estimates effective compute grows at least 10× annually, with algorithmic efficiency alone improving at 6× per year.

If this growth maintains, five years from now effective compute will be 100,000 times greater. An AGI costing US$1 million per instantiation today becomes practical for mass deployment within years. The investment backing this growth is staggering: global AI capital expenditure is projected to reach US$5.5 trillion through 2030, with hyperscaler spending exceeding US$1.1 trillion in 2027 alone. In 2026, global AI investment will exceed US$1 trillion for the first time.

Model Scale: The Leading Edge

Recent estimates place Anthropic's Mythos 5 at approximately 8 trillion parameters — a scale now being matched by ByteDance, which is reportedly training a 10 trillion-parameter model. This represents a significant leap from previous benchmarks: DeepSeek V4-Pro sits at 1.6 trillion, while Moonshot AI's Kimi K3 has 2.8 trillion.

Yet as experts note, direct comparisons remain difficult because companies increasingly do not disclose parameter counts. The opacity itself is a signal — labs stopped publishing because the number stopped meaning what it used to.

Performance Benchmarks

The real signal now lies in capability scores. Gemini 3.1 Pro reportedly achieved 100% on AIME 2025 with code execution enabled — a first in the field. DeepSeek V3 scored 93.1%, placing it above o3-mini and in the same tier as Grok-3 at 93.3%.

Google DeepMind's projections suggest that 1 billion AGI instances sharing experience and thinking at accelerated speed would constitute a superintelligence, even if individual models never exceed human baseline capabilities.

Timelines and Safety: The Governance Numbers

Expert forecasts have shifted significantly toward near-term AGI. Median forecasts place AGI by 2050, with a 35% probability assigned to AI matching electricity in societal impact by 2040 . Yet the allocation of resources is deeply uneven: estimates suggest 20,000 people work on AGI development, while only about 200 focus on safety.

For Global Future Nexus, these numbers are not abstract statistics. They define the terrain of responsible integration: the velocity of progress, the scale of investment, and the governance gaps that demand attention. The parameter count alone no longer tells the story. What matters is how we channel the intelligence those parameters enable toward planetary sustainability and human flourishing.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Previous
Previous

AGI and the future of cyber-physical systems

Next
Next

The velocity gap: when AI creates faster than humans can verify