China’s Kimi K3 ‘significantly below’ US rivals in hacking power, study shows
A joint British-US government study found that China's Kimi K3 large language model performs significantly worse than top US rivals in its ability to launch cyberattacks, challenging US anxieties about Chinese AI.
Intelligence analysis by Gemini 2.5 Flash

The research, conducted by the UK Artificial Intelligence Security Institute and the US Centre for AI Standards and Innovation, used the ExploitBench benchmark to assess AI models. Kimi K3 scored 32.2 per cent, far behind unnamed US models averaging 76.2 per cent, and failed to achieve high-level exploits, suggesting a considerable gap in advanced cyber capabilities.
Imagine a super-smart computer program that can find secret ways to break into other computers, like a digital detective. A new study looked at China's best program, Kimi K3, and found it's not as good at finding these secret ways as the top programs from the US. It's like Kimi K3 can only pick simple locks, while the US programs can pick really complicated ones and even get inside the whole house. This means that for now, the US programs are much better at digital 'hacking' than China's Kimi K3.
Analysis
Assessing Cyberattack Capabilities
A recent joint study by the UK Artificial Intelligence Security Institute (AISI) and the US Centre for AI Standards and Innovation (CAISI) has shed light on the cyberattack capabilities of China's Kimi K3 large language model. The research utilized ExploitBench, a public benchmark designed to evaluate an AI's proficiency in developing cybersecurity exploits. Kimi K3, developed by Chinese unicorn Moonshot AI and considered one of China's most powerful LLMs, achieved an overall score of 32.2 per cent.
This performance, while outperforming domestic rival Zhipu AI’s GLM-5.2 (24.4 per cent), was significantly lower than the average of 76.2 per cent recorded by top, unnamed US models. A critical finding was Kimi K3's inability to achieve arbitrary code execution—the highest level of exploit that grants full control over a target system—across any of the 41 ExploitBench tasks. In contrast, leading US models successfully achieved this on 20 tasks, highlighting a substantial disparity in advanced hacking power.
Implications for the AI Tech Race
The findings of this study directly address Washington's growing concerns regarding the rapid advancement of Chinese open-source artificial intelligence. The report suggests that, at least in the domain of cyberattack capabilities, Chinese models like Kimi K3 are not yet on par with their leading American counterparts. This could potentially temper some of the immediate anxieties in the US about China's ability to leverage AI for sophisticated cyber warfare or espionage.
However, the study also underscores the dynamic nature of the AI tech race. While Kimi K3 currently lags, the pace of AI development is incredibly fast, and capabilities can evolve quickly. The report provides a snapshot in time, and continued investment and research in China could narrow this gap in the future. It also highlights the importance of transparent benchmarking and international collaboration in assessing AI risks and capabilities.
The Broader AI Security Landscape
This research contributes to a broader global effort to understand and mitigate the security implications of advanced AI models. As AI becomes more powerful, its potential for misuse in cyberattacks, disinformation, and other malicious activities grows. Studies like this are crucial for policymakers and security experts to develop appropriate safeguards, regulations, and international standards.
By publicly assessing the capabilities of frontier models, the AISI and CAISI are promoting a more informed discussion about AI safety and security. The disparity observed between Chinese and US models could influence strategic decisions regarding AI development, export controls, and defensive cybersecurity measures. Ultimately, a clear understanding of these capabilities is essential for fostering responsible AI innovation and preventing unintended consequences in an increasingly interconnected and digitally vulnerable world.
Key points
- China's Kimi K3 large language model scored 32.2 per cent on ExploitBench, a benchmark for AI cyberattack capabilities.
- Top US models averaged 76.2 per cent on the same benchmark, significantly outperforming Kimi K3.
- Kimi K3 failed to achieve arbitrary code execution, the highest-level exploit, in any of the 41 tasks.
- The study was conducted by the UK Artificial Intelligence Security Institute and the US Centre for AI Standards and Innovation.
- The findings challenge US anxieties about the rapid rise of Chinese open-source AI in the cybersecurity domain.
The study's findings could alleviate some immediate US concerns regarding China's AI-driven cyberattack capabilities, potentially fostering a more measured approach to AI policy and competition. It might also encourage Chinese developers to prioritize robust security measures and ethical AI development, leading to safer global AI systems.
The significant gap in cyberattack capabilities could intensify the global AI arms race, with China potentially redoubling efforts to catch up, leading to increased investment in offensive AI research. This disparity also highlights the ongoing challenge of accurately assessing and mitigating AI-driven cyber risks, potentially leaving some nations more vulnerable.
