16 Zhang B200 Could Run Kimi K3, 8 AMD Cards Can Fit Instead
A recent test by Wafer AI on AMD MI355X showed that 16 NVIDIA B200 cards could be replaced by 8 AMD cards, achieving a 3.8 times higher throughput and a 1.7 times higher single-user generation speed. The test also highlighted the importance of memory capacity in large-sca…
Intelligence analysis by Llama

Wafer AI's test on AMD MI355X showed that 8 cards can replace 16 NVIDIA B200 cards, achieving a 3.8 times higher throughput and a 1.7 times higher single-user generation speed. The test also highlighted the importance of memory capacity in large-scale model deployment.
Imagine you have a big puzzle with millions of pieces. You need a lot of memory to store all the pieces, so you can solve the puzzle. Wafer AI tested how many pieces of the puzzle they could fit into a computer with 8 AMD cards. They found that they could fit more pieces than if they used 16 NVIDIA cards. This is important because it means that AMD cards can be used to solve bigger puzzles, which is useful for things like language translation and image recognition.
Analysis
A $60B Vote of Confidence
Wafer AI's recent test on AMD MI355X has sent shockwaves through the AI community, as it showed that 16 NVIDIA B200 cards could be replaced by 8 AMD cards, achieving a 3.8 times higher throughput and a 1.7 times higher single-user generation speed. The test also highlighted the importance of memory capacity in large-scale model deployment.
The test was conducted on a model with 2.8 billion parameters, which required over 1.5 TB of memory to run. The Wafer AI team used 8 AMD MI355X cards, each with 288 GB of memory, to run the model, achieving a peak throughput of 952 Token/s and a single-user generation speed of 118 Token/s. In contrast, the same model running on 16 NVIDIA B200 cards, each with 192 GB of memory, achieved a peak throughput of 498 Token/s and a single-user generation speed of 90 Token/s.
The results of the test have significant implications for the deployment of large-scale models in data centers. As models continue to grow in size, memory capacity becomes increasingly important. The test shows that AMD cards can provide a significant advantage in terms of memory capacity, making them an attractive option for data centers looking to deploy large-scale models.
Why Cursor?
Wafer AI's test on AMD MI355X also highlights the importance of software support for large-scale model deployment. The team was able to run the model on the AMD cards without any issues, thanks to the support provided by the ROCm team. This is a significant advantage over NVIDIA cards, which often require significant software modifications to run large-scale models.
The Road Ahead
The results of Wafer AI's test on AMD MI355X have significant implications for the future of large-scale model deployment. As models continue to grow in size, memory capacity becomes increasingly important. The test shows that AMD cards can provide a significant advantage in terms of memory capacity, making them an attractive option for data centers looking to deploy large-scale models. Additionally, the test highlights the importance of software support for large-scale model deployment, and the need for data centers to invest in software that can support large-scale models.
Key points
- Wafer AI's test on AMD MI355X showed that 16 NVIDIA B200 cards could be replaced by 8 AMD cards, achieving a 3.8 times higher throughput and a 1.7 times higher single-user generation speed.
- The test highlighted the importance of memory capacity in large-scale model deployment.
- Wafer AI's test on AMD MI355X also showed that the team was able to run the model on the AMD cards without any issues, thanks to the support provided by the ROCm team.
The results of Wafer AI's test on AMD MI355X suggest that AMD cards may be a viable option for large-scale model deployment in the future. If AMD can continue to improve the performance and memory capacity of their cards, they may be able to compete with NVIDIA in the market for large-scale model deployment.
However, it's worth noting that the test was conducted on a specific model and configuration, and it's unclear whether the results would generalize to other models and configurations. Additionally, the test was conducted on a relatively small scale, and it's unclear whether the results would scale up to larger models and configurations.
Market signals
- XAU Escalation drives safe-haven demand for gold, per the article's framing of investor reaction.
- OIL Supply-route risk from the reported conflict pushes oil prices higher.
AI-generated analysis of potential market relevance. Not financial advice.


