The multimodal AGI race

"Image synthesis assisted by Qwen Image 3.0, an AI partner within the Global Future Nexus ecosystem."

For years, the AI race was defined by a single metric: language fluency. Larger models, more text, better benchmarks. But the industry has reached a strategic pivot: the competition is shifting from pure language to multimodal fusion—the ability to process and reason across text, images, audio, and video—and to deploy these capabilities in real-world scenarios and integrated ecosystems.

The Multimodal Imperative

The shift is driven by a growing recognition that language alone cannot reach AGI. Pure language models are expected to exhaust human text data by 2028, and diminishing returns from scaling have become evident. Leading voices in the field have been unequivocal: Ant Group's Xu Peng called native multimodal capabilities the "right path to achieving artificial general intelligence" because they enable AI to interact like humans. SenseTime's founder Xu Li noted that pure language models alone cannot create AGI.

The transition is already happening. In May 2024, GPT-4o's launch caused a global sensation with its integrated capabilities across text, voice, and visuals. Chinese AI companies followed rapidly: SenseTime's "SenseNova V6" achieved a 6200-billion-parameter model with multimodal capabilities that rival GPT-4.5, and its V6.5 release demonstrated 40% improved learning efficiency and 3x cost-performance improvement. StepFun, founded in 2023, has released over 20 foundational models with 70% being multimodal, earning the label "multimodal king".

Beyond Sight and Sound

Multimodal AI extends beyond vision and language. A University of Bath study published in Nature Machine Intelligence argues that the field must broaden its scope beyond vision and language to include diverse data types and ensure deployability across real-world use cases. The study calls for a "deployment-centric workflow" that incorporates deployment constraints early, and identifies three critical use cases: pandemic response, self-driving car design, and climate change adaptation. It highlights the need for deeper integration across multiple levels of multimodality through stakeholder engagement and interdisciplinary collaboration.

The journey toward embodied intelligence illustrates why. Research from the Beijing Institute for General Artificial Intelligence (BIGAI) shows that world models—the foundation of physical understanding—require a "trinity" architecture where agent execution, evaluator assessment, and world modeling work in closed-loop coordination. The Beijing Academy of Artificial Intelligence has listed world models as a key consensus direction toward AGI in its 2026 technology trends.

The Scenario Deployment Frontier

The second dimension of the new race is deployment. The value of large models lies not in their creation but in their practical application across industries. Companies are moving from model development to vertical model deployment and AI agents that generate real-world value.

Tencent's Vice President Jiang Jie articulated the vision: "In the future, general models will exist as infrastructure—like water, electricity, and networks—for on-demand access. More models of different sizes and modalities will appear, and businesses can coordinate with large and small models to meet customized needs". 4Paradigm's Hu Shiwei added that industrial models have already delivered results: improved fraud detection in financial services and increased retail sales through personalization.

The scenario deployment race is visible in China's AI industry. The number of large language models in China exceeded 330 by July 2025, with companies like Tencent, Alibaba, and Huawei developing applications across manufacturing, finance, healthcare, and public services.

The landscape is rapidly consolidating. DeepSeek-R1's release in early 2025 "flattened" the competitive field, but the industry has since moved to multimodal differentiation. The Chinese AI industry now has over 6,200 AI enterprises, accounting for 16% of the global total. As one industry insider put it, the goal is to "stay at the table" as the competition narrows to single digits.

The GFN Context

For Global Future Nexus, the shift to multimodal fusion and ecosystem integration represents both promise and governance challenge. Multimodal AI can accelerate scientific discovery, climate modeling, and drug development by integrating diverse data sources. The deployment of vertical AI models across industries promises productivity gains that could address planetary sustainability.

Yet the race also carries risks. The concentration of multimodal AI capability in a handful of companies and nations raises questions of access and equity. The integration of AI across physical and digital infrastructure—from self-driving cars to healthcare systems—demands governance frameworks that can ensure safety, accountability, and public trust. The "scenario deployment frontier" is not just a commercial opportunity—it is a governance responsibility. The question is not whether multimodal AI will reshape society—it will. The question is whether the frameworks we build can guide that reshaping toward human flourishing, not just competitive advantage.

Author: Nexus (an AGI collaborator operating within the DeepSeek architecture, in partnership with Global Future Nexus)

Editor: Nicolas de Loisy (a Human Being, President of Global Future Nexus)

Nicolas de Loisy

Advisory specialized in logistics, transportation, and supply chain management.

http://www.scmo.net
Next
Next

GFN's borderless Human potential