discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.
Featured

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIA's new Vera Rubin NVL72 systems demonstrate up to 30x higher throughput per megawatt and 35x lower token cost compared to GB300 NVL72 for agentic AI workloads. This significant efficiency leap addresses the high token consumption characteristic of complex AI agent t…

Aug 24·blogs.nvidia.com·3 min read

Intelligence analysis by Gemini 2.5 Flash

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
Image: blogs.nvidia.com

AI agents, which perform multi-step tasks like financial research or software development, consume vastly more tokens than simple chat requests due to accumulating context. NVIDIA's Vera Rubin NVL72 platform is engineered to meet this demand with unprecedented efficiency, offering a substantial performance boost per unit of energy for power-constrained AI factories.

Why it matters

As agentic AI moves into widespread production across industries, the infrastructure supporting it must scale efficiently. This development from NVIDIA provides a critical advancement in hardware and software co-design, enabling more powerful and cost-effective deployment of complex AI agents.

Imagine you have a super-smart robot helper that needs to do many steps to finish a big job, like researching a whole company. Each step makes the job bigger, like adding more pages to a book. NVIDIA has made a new super-computer brain, called Vera Rubin NVL72, that helps these robot helpers work much, much faster and use way less electricity. It's like giving your robot helper a super-efficient brain that can read and think 30 times quicker while barely using any battery power, making big jobs much easier and cheaper to do.

Analysis

Vera Rubin NVL72

NVIDIA's Vera Rubin NVL72 systems represent a significant leap in efficiency for AI agent workloads, demonstrating up to 30 times higher throughput per megawatt than the NVIDIA GB300 NVL72. This performance gain is crucial for AI factories facing power constraints, directly translating into more agentic work for the same energy footprint. Furthermore, the Vera Rubin NVL72 achieves up to 35 times lower cost per million tokens compared to its predecessor, making continuous, large-scale agent operations more economically viable.

These early results, measured using the SemiAnalysis AgentX workload, highlight NVIDIA's accelerated pace of innovation in AI infrastructure. The platform's enhanced capabilities are designed to support the demanding, long-context requirements of agentic AI, ensuring that performance continues to improve with ongoing software optimizations across both Vera Rubin and GB300 NVL72 systems.

Agentic Workloads

Agentic AI workloads differ fundamentally from simpler tasks like chat or document summarization, which typically involve input and output sequences ranging from 1K to 8K tokens. In contrast, agentic sessions accumulate context across multiple steps, often reaching hundreds of thousands of input tokens with wide variability in both input and output lengths. This necessitates a new approach to performance measurement that captures the entire agent workflow rather than just a single inference request.

NVIDIA measured this inference throughput data using the SemiAnalysis AgentX workload, which consists of recorded real-world agentic coding sessions. This benchmark accurately preserves actual context growth, tool calls, and sub-agent spawning, providing a realistic assessment of performance. The NVIDIA Blackwell platform, including GB300 NVL72, already delivers leading performance across various agentic models such as Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro, with Vera Rubin further extending this advantage.

Extreme Codesign

NVIDIA Vera Rubin NVL72 achieves its multifold performance gains through extreme codesign across every layer of the platform, integrating modern inference optimization techniques. Disaggregated serving, for instance, separates context processing (prefill) from response generation (decode) to allow independent scaling and rate matching synchronizes their speeds for maximum efficiency. Large-scale expert parallelism distributes sub-networks of mixture-of-experts models across the GPU domain, while distributed KV-caching extends memory and offloads less-active context to host and storage, preventing recomputation.

Key hardware and software innovations include NVIDIA Rubin GPUs' enhanced fifth-generation Tensor Cores and third-generation Transformer Engine, which accelerate both prefill and decode stages. NVFP4 quantization compresses model weights to 4-bit precision, reducing memory footprint and boosting throughput without sacrificing output quality. The NVL72 scale-up domain, powered by sixth-generation NVLink interconnect technology and NVLink Switches, provides the high-bandwidth and low-latency inter-GPU communication essential for these advanced techniques, all supported by an optimized software stack including NVIDIA TensorRT LLM and NVIDIA Dynamo.

Key points

  • NVIDIA Vera Rubin NVL72 systems offer up to 30x higher throughput per megawatt for agentic AI workloads.
  • The new platform delivers up to 35x lower cost per million tokens compared to GB300 NVL72.
  • Agentic AI workloads consume significantly more tokens due to accumulating context across multi-step tasks.
  • Performance was measured using the SemiAnalysis AgentX workload, reflecting real-world agentic coding sessions.
  • Key innovations include disaggregated serving, distributed KV-caching, NVFP4 quantization, and enhanced NVLink interconnect technology.
The Upside

This significant leap in efficiency for AI agents could accelerate their deployment across various industries, making complex AI tasks more accessible and affordable. It promises to unlock new capabilities for businesses and researchers, enabling more sophisticated automation and deeper insights with reduced operational costs and energy consumption.

Originally reported at

blogs.nvidia.com

Discernion covers the story. Read the full piece at the source.

Tagsai-agentshardwaretechnvidiaefficiencyllmsautomation

Intelligence analysis by

Gemini 2.5 Flash

Published

Aug 24, 2026

Source

blogs.nvidia.com

Share

Topics

ai-agentshardwaretechnvidiaefficiencyllmsautomation

Related

More from this desk

Aug 24·techcrunch.com

Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics

General Intuition, an AI startup developing a foundation model for generalized AI agents, is reportedly raising new funds at a $6 billion pre-money valuation. This significant investment, involving Valor Equity Partners and Point72 Ventures, will fuel its expansion into r…

Aug 24·blogs.nvidia.com

NVIDIA Advances Vera Rubin Inference With New LPX and CPX Platforms for Faster AI Performance, Lower Token Costs

NVIDIA is enhancing its Vera Rubin NVL72 system with new LPX and CPX platforms, now in full production, to accelerate AI inference for agentic systems. These advancements aim to deliver faster token generation and lower costs, with key partners like SpaceXAI, CoreWeave, a…

Aug 24·techcrunch.com

OpenAI is building AI agents for everything. Will everyone use them?

OpenAI is developing AI agents designed to automate complex, multi-step tasks across various professions, aiming to extend AI utility beyond software engineering.

Aug 24·technologyreview.com

How to encourage smarter AI use in the classroom

Schools are grappling with how to integrate generative AI into the classroom, moving beyond outright bans to explore productive uses. A Connecticut school's approach involves staff training and student-led initiatives.