NVIDIA Blackwell Leads on First Agentic AI Infrastructure Benchmark
NVIDIA says its Blackwell Ultra NVL72 tops AgentPerf, a new benchmark for agentic AI systems, with up to 20x more agents per megawatt than Hopper.
Intelligence analysis by GPT-5.4 Mini

NVIDIA is touting early AgentPerf results as a new way to judge infrastructure for agentic AI, where many LLM and tool calls are chained together. The company says Blackwell Ultra and GB300 NVL72 lead the first published results and that partners are already using Blackwell for production agent workloads.
NVIDIA says it made a new test for AI workers that do many steps, not just one answer. It claims its newest chips are best at this and can do far more AI jobs using the same power, like fitting more delivery vans into the same garage.
Analysis
What NVIDIA is claiming
NVIDIA says Artificial Analysis’ new AgentPerf benchmark is the first benchmark focused on agentic AI infrastructure. Unlike ordinary inference tests that measure one model response at a time, AgentPerf is meant to reflect workloads where an agent reads files, calls tools, writes code, runs commands, and iterates over many steps.
Why the benchmark is different
The article argues that agentic AI is not just more chat. A single task can involve dozens or even hundreds of LLM calls plus tool calls, with growing context and delays at each handoff. NVIDIA says this makes performance measurement much harder, because the real question is not only how fast one model responds, but how many useful agent tasks a system can support at once while meeting responsiveness targets.
The published results
In the first round of results, NVIDIA says Blackwell Ultra NVL72 delivers leading performance across the workloads tested, running 20x more agents per megawatt than NVIDIA Hopper. For the DeepSeek V4 Pro workload, NVIDIA says GB300 NVL72 reaches the highest benchmark performance and can run up to 20x more agents per megawatt than HGX H200, including at service-level objectives of 20 and 60 tokens per second per agent.
NVIDIA attributes the gains to full-stack design: 72 GPUs in one rack-scale system, CUDA kernels that overlap communication with compute, and TensorRT LLM optimizations that separate input processing from output generation.
How AgentPerf was built
The benchmark uses real coding-agent trajectories drawn from public repositories across more than 12 programming languages. Tool calls are simulated rather than executed, so the results are intended to reflect accelerated computing performance rather than software execution variability. NVIDIA says the benchmark measures how many agentic tasks a platform can support simultaneously while still meeting its thresholds.
Ecosystem adoption
The company also says providers including Baseten, DeepInfra, and Together AI are already serving agentic workloads on Blackwell. It cites Together AI’s work powering Cursor and DeepInfra’s support for Pam.ai. NVIDIA closes by pointing to Vera Rubin, which it says is now in full production.
Key points
- AgentPerf is presented as the first benchmark built specifically for agentic AI infrastructure.
- NVIDIA says Blackwell Ultra NVL72 leads the first published results, with up to 20x more agents per megawatt than Hopper.
- The benchmark is based on real coding-agent trajectories from public repositories across 12+ programming languages.
- NVIDIA says its gains come from rack-scale GPU design, CUDA overlap of compute and communication, and TensorRT LLM optimizations.
- The company says Baseten, DeepInfra, and Together AI are already running agentic workloads on Blackwell.
If AgentPerf becomes a common yardstick, buyers may get a better way to compare systems for real agent workloads instead of chat-style tests. That could help teams choose infrastructure that runs more agents per watt and per dollar, which is exactly what the article says enterprises need. The ecosystem notes also suggest Blackwell is already moving from benchmark claims into production use.
The benchmark is new, so its influence will depend on whether the industry accepts its methodology as representative of production agent work. Because the first results come from NVIDIA's own blog and highlight NVIDIA systems, readers may want more independent validation before treating the performance claims as settled. Simulated tool calls also mean the benchmark does not measure every part of a live agent system.



