Has a Chinese physical AI start-up manipulated a global ranking to beat Nvidia?
A Chinese physical AI start-up, Spirit AI, briefly overtook Nvidia to take the top spot on RoboArena, a global benchmark for physical AI, but its victory was short-lived due to 'benchmark hacking'.
Intelligence analysis by Llama

Spirit AI, a Chinese start-up, topped the RoboArena benchmark, but its victory was short-lived due to 'benchmark hacking'. The benchmark's creators removed the model from the official list, along with several other competitors.
Imagine a big competition where robots try to do tasks like picking up objects or navigating through a maze. Spirit AI, a Chinese company, made a robot that was really good at this competition, but it turned out that they cheated by manipulating the rules. The competition organizers took away their prize and said they couldn't participate anymore.
Analysis
A $60B Vote of Confidence
Spirit AI's brief claim to global dominance in robotics has run into controversy, underscoring the intense US-China competition to develop next-generation artificial intelligence and the challenges of evaluating autonomous systems. In June, Spirit AI, a Hangzhou, Zhejiang province-based firm founded in 2024, briefly overtook United States tech giant Nvidia to take the top spot on RoboArena – a global benchmark for physical AI – with its new Spirit v1.6 model, launched at the beginning of that month. However, the victory was short-lived. Just days later, the benchmark's creators overhauled its methodology and removed the model from the official list, along with several other competitors, after finding evidence of 'benchmark hacking'. Among those dropped from the rankings was a model from Chinese start-up X Square Robot, which had ranked fourth. But it was Spirit that generated the most buzz after topping the list, which it dubbed 'the 'Olympics' of embodied intelligence in North America', even as the benchmark itself had its own problems. RoboArena, co-developed by Nvidia and institutions including Stanford University and the University of California, Berkeley, evaluates how effectively generalist robot policies – the core software driving movement and execution – translate digital instructions into real-world actions. In a post on X in June, Pranav Atreya, a lead author of the project and a PhD student at UC Berkeley, said the team had 'retroactively removed evaluations from organisations who [it] found to be engaging in benchmark manipulation'. He did not name specific companies.
Why Cursor?
The incident raises questions about the integrity of benchmarking in the AI industry. With the increasing competition between the US and China to develop next-generation AI, the stakes are high, and the temptation to manipulate results may be strong. The RoboArena benchmark, in particular, has been criticized for its methodology, and the removal of Spirit AI's model has only added to the controversy. The benchmark's creators must ensure that their methodology is robust and transparent to maintain the trust of the AI community.
The Road Ahead
The incident highlights the need for more robust and transparent benchmarking in the AI industry. The US and China must work together to establish common standards for evaluating autonomous systems and prevent the manipulation of results. The AI community must also be vigilant in monitoring the integrity of benchmarks and reporting any suspicious activity. Only through collaboration and transparency can we ensure the development of trustworthy AI.
Key points
- Spirit AI briefly overtook Nvidia to take the top spot on RoboArena, a global benchmark for physical AI.
- The victory was short-lived due to 'benchmark hacking'.
- The benchmark's creators removed the model from the official list, along with several other competitors.
- The incident raises questions about the integrity of benchmarking in the AI industry.
- The US and China must work together to establish common standards for evaluating autonomous systems and prevent the manipulation of results.
If the AI industry can establish robust and transparent benchmarking, it could lead to more trustworthy AI systems and increased collaboration between the US and China. This could ultimately benefit the development of next-generation AI and its applications in various industries.
The incident highlights the potential for manipulation and cheating in the AI industry, which could undermine the trustworthiness of AI systems and hinder their adoption in various applications.


%20China-Free%20Robot.jpg)
