Alibaba’s new AI model scores higher than OpenAI, Google rivals in coding ranking
Alibaba’s Qwen3.7-Max ranked fourth on Code Arena, ahead of OpenAI and Google models.
Intelligence analysis by GPT-5.4 Mini

Alibaba’s latest model, Qwen3.7-Max, reached fourth place on Code Arena with a score of 1,541. It was the only non-US model in the top five, which were otherwise dominated by Anthropic’s Claude variants.
Alibaba made a new robot brain called Qwen3.7-Max. On a big coding contest board, it came in fourth place.
This contest is like a bake-off for computer helpers. Instead of answering easy quiz questions, the models have to build real web apps, and people vote on which ones work best.
The important part is that only one company from outside the United States made the top five. That shows the race to build better coding helpers is getting tighter.
Analysis
What happened
Alibaba Group Holding’s latest AI model, Qwen3.7-Max, placed fourth on the Code Arena coding leaderboard with a score of 1,541. According to the article, that put it ahead of rival models from OpenAI and Google and made Alibaba the only non-US developer in the top five.
Why this leaderboard matters
The story says Code Arena differs from older coding benchmarks like HumanEval and SWE-bench because it tests whether models can independently build complete, interactive web applications from scratch based on user prompts. The ranking is also shaped by blind user voting on anonymized outputs, which the article frames as a better reflection of what real-world developers prefer.
Bigger picture
The article ties the result to a broader shift among Chinese AI developers away from general-purpose chatbots and toward coding agents and other autonomous systems. That matters because investors increasingly see these tools as among the most commercially viable uses for generative AI. The top five on this leaderboard were otherwise filled by Anthropic’s Claude models, which underscores how concentrated the frontier coding race remains even as Alibaba closes the gap in this specific benchmark.
Key points
- Alibaba’s Qwen3.7-Max ranked fourth on Code Arena with a score of 1,541.
- It was the only non-US model in the top five.
- The top five were otherwise made up of Anthropic’s Claude models.
- Code Arena tests whether models can build complete interactive web apps from scratch.
- The article says Chinese developers are shifting toward coding agents and autonomous systems.



