The Native Multimodal Time Gap Between Kimi K3 and DeepSeek V4
Chinese large model developers are diverging on integrating native multimodal capabilities, particularly vision, into their foundational models, with Kimi K3 embracing it while DeepSeek V4-Flash prioritizes text-based optimization.
Intelligence analysis by Gemini 2.5 Flash

The article highlights a strategic split among leading Chinese AI companies regarding the timing and necessity of native multimodal integration. While some, like Moonshot AI's Kimi K3, are heavily investing in vision-in-the-loop for advanced Agent tasks, others, such as DeepSeek, are focusing on refining text and coding abilities through post-training, acknowledging multimodal's long-…
Imagine you're building with LEGOs, and you have a robot helper. Some robots can only read instructions written down, so if a piece is in the wrong spot, you have to tell them in words. But a robot with 'eyes' (multimodal ability) can look at the LEGO model, see if a piece is crooked, and fix it all by itself, making the building much smoother and faster. Chinese AI companies are debating if their robot helpers should get 'eyes' now or later, because adding eyes makes them smarter but also takes a lot more effort and special training.
Analysis
The Diverging Paths of Chinese LLMs
The landscape of Chinese large language model development is currently marked by a significant strategic divergence, particularly concerning the integration of native multimodal capabilities. On one side, companies like Moonshot AI, with their Kimi K3 model, are aggressively pursuing a 'large and comprehensive' approach, embedding vision directly into the foundational model's pre-training. This strategy aims to equip AI Agents with 'eyes' from the ground up, enabling them to perceive and interpret visual information directly, which is crucial for complex tasks like web development where visual feedback is paramount.
Conversely, DeepSeek, with its V4-Flash model, represents a more text-centric approach. While acknowledging the long-term importance of multimodal capabilities, DeepSeek's founder initially suggested that advanced AI doesn't immediately require world models or even multimodality. Their current focus is on optimizing core text, coding, and Agent abilities through sophisticated post-training methods, without a substantial increase in parameters or immediate native vision integration. This reflects a belief that significant performance gains can still be achieved within the text domain, offering a more resource-efficient path in the short term.
Why Agents Need "Eyes": The Multimodal Advantage
The article underscores the growing necessity for AI Agents to possess visual understanding, especially as their task chains become longer and more intricate. For tasks like generating interactive web pages, a model's ability to 'see' the output and identify visual discrepancies (e.g., layout errors, incorrect styling) is invaluable. Kimi K3's 'vision in the loop' mechanism exemplifies this, allowing the model to iteratively refine code based on visual feedback, leading to more accurate and user-aligned results.
Proponents of native multimodal integration argue that it offers a superior 'communication channel' between visual input and the language core, unlike modular solutions that convert images to text before processing. This direct integration allows for a more holistic understanding and faster, more accurate decision-making in scenarios requiring continuous visual assessment. The experience of GPT-5.6 Sol operating a PC Agent, handling complex tasks like browser interaction and code deployment, highlights the transformative potential when models can 'see' and adapt to their environment.
The Cost and Timing of Vision Integration
Despite the clear advantages, integrating native multimodal capabilities comes with substantial trade-offs. It demands significantly more computational resources, data, and model capacity compared to purely text-based models. The article notes that adding visual data can initially degrade existing language, reasoning, and coding abilities, requiring careful balancing and extensive fine-tuning. The '1+1 is greater than 2' analogy for resource expenditure emphasizes the complexity of training a unified multimodal model.
Companies are grappling with the question of how much resource allocation to dedicate to vision at this stage, especially when rapid advancements are still being made in text and coding. The decision is not purely technical but also influenced by leadership's research background, product strategies, and training costs. While coding and Agent capabilities are seen as the immediate 'entry ticket' to market competition and commercialization, the long-term evolution towards a more comprehensive understanding of the world, as advocated by figures like Yann LeCun, points towards the eventual necessity of multimodal learning. This creates a dual timeline for Chinese LLM developers: one for immediate competitive gains and another for long-term foundational advancements.
Key points
- Chinese AI companies are divided on the immediate integration of native multimodal capabilities into large language models.
- Kimi K3 from Moonshot AI emphasizes native multimodal vision for advanced Agent tasks, using 'vision in the loop' for error correction.
- DeepSeek V4-Flash prioritizes text, coding, and Agent performance through post-training, deferring extensive native multimodal integration.
- Native multimodal models offer deeper integration and better feedback for complex tasks but require significantly more resources and can initially impact other capabilities.
- The industry consensus is that multimodal is important long-term, but companies differ on the timing and cost of its implementation, influenced by various strategic factors.
The push towards native multimodal capabilities, as seen with Kimi K3, could lead to significantly more capable and intuitive AI Agents that can understand and interact with the digital and real world more effectively. This could unlock new applications in areas like web development, design, and complex automation, making AI tools more versatile and user-friendly.
The high resource demands and potential for initial performance degradation when integrating native multimodal capabilities could slow down the development of other crucial AI functionalities like coding and reasoning. Companies that prioritize this route might face higher costs and a longer path to commercialization compared to those focusing on optimizing text-based models.



