discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.
Featured

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

NVIDIA has launched its Vera Rubin platform, an AI supercomputer designed for gigascale operations, emphasizing extreme co-design for superior performance per watt and reduced token costs.

Jul 21·blogs.nvidia.com·3 min read

Intelligence analysis by Gemini 2.5 Flash

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
Image: blogs.nvidia.com

The Vera Rubin NVL72 system is ramping up production with major cloud partners, featuring a custom-designed CPU, advanced networking, and liquid cooling to accelerate AI factories and support the growing demand for efficient, high-throughput AI compute.

Why it matters

This development is crucial for the AI industry as it promises significantly lower operational costs and higher efficiency for large-scale AI model training and inference, directly addressing power constraints and accelerating the deployment of advanced AI applications globally.

Imagine a super-fast brain for computers that helps them learn and think, like a super-smart robot. NVIDIA made a new version called Vera Rubin that's like a super-efficient brain. It uses much less electricity and costs less to run, so big companies can build more of these smart computer brains without using too much power or spending too much money. It's like getting a super-powered toy that uses very little battery.

Analysis

Engineering for Gigascale Efficiency

NVIDIA's Vera Rubin platform represents a significant leap in AI infrastructure, built on an 'extreme co-design' philosophy that integrates seven chips and five rack trays into a single, unified system. At its core is the NVIDIA Vera CPU, featuring custom Olympus cores that deliver twice the single-threaded performance and three times the core-to-core bandwidth compared to competing chiplet designs. This meticulous engineering aims to optimize for 'agentic workloads,' which are increasingly critical for advanced AI applications. The platform also incorporates sixth-generation NVLink for scale-up, offering over double the throughput and three times lower latency, alongside Spectrum-X Ethernet for scale-out, which provides 1.6x higher RDMA bandwidth. These networking advancements, coupled with NVIDIA Photonics' co-packaged optics, significantly reduce power consumption and improve reliability, making the Vera Rubin system a highly integrated and efficient solution for demanding AI tasks.

Powering Global AI Factories and Sovereign AI

The Vera Rubin platform is rapidly being adopted by leading AI infrastructure builders and cloud providers, including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, SpaceXAI, and Tesla. This widespread adoption underscores the industry's need for high-performance, scalable AI compute. A notable partnership highlighted in the article is with Microsoft and Mistral in Europe, where Vera Rubin will underpin a multibillion-dollar agreement to expand AI infrastructure. This initiative is particularly focused on enabling 'sovereign-ready AI,' allowing European governments and regulated industries to deploy AI solutions that adhere to regional data control, governance, and autonomy requirements. By providing a robust computing foundation, Vera Rubin is positioned to accelerate Europe's open-model ecosystem and support the deployment of frontier AI models like Mistral Medium 3.5 and OCR 4 within secure, localized environments.

The Economic Impact of Token Cost Reduction

One of the most compelling aspects of the Vera Rubin platform is its promise of dramatically lower token costs and superior performance per watt. Benchmarks, such as CoreWeave's DeepSeek-R1 test, indicate a tenfold increase in throughput per megawatt compared to the Grace Blackwell NVL72, translating to one-tenth the cost per million tokens. This efficiency gain is critical for 'power-constrained AI factories' and for handling the escalating demands of agentic systems, which can consume up to 15 times more tokens than traditional AI applications. Beyond raw performance, the system's design also addresses operational efficiencies, such as reducing compute tray assembly time from hours to minutes and enabling chiller-free dry-cooler operation, saving millions of gallons of water annually. These combined innovations offer substantial economic advantages, making advanced AI more accessible and sustainable for partners worldwide.

Key points

  • NVIDIA's Vera Rubin platform is a gigascale AI supercomputer designed for extreme efficiency and performance.
  • It features a custom Vera CPU, advanced NVLink and Spectrum-X networking, and innovative liquid cooling.
  • The platform delivers up to 10x more throughput per megawatt and one-tenth the cost per million tokens compared to previous generations.
  • Major cloud providers and AI companies like CoreWeave, Microsoft, Google Cloud, and Oracle are adopting Vera Rubin.
  • Vera Rubin is central to expanding AI infrastructure in Europe, supporting 'sovereign-ready AI' initiatives with partners like Mistral.
The Upside

The Vera Rubin platform is expected to significantly accelerate AI development and deployment by offering unprecedented performance per watt and lower token costs. This could lead to more powerful and accessible AI models, fostering innovation across various industries and enabling the creation of advanced agentic systems.

Market signals

NVDA· NASDAQ
  • NVDA The launch of the Vera Rubin platform with strong performance metrics and major partner adoption signals continued leadership and growth for NVIDIA in the AI hardware market.

AI-generated analysis of potential market relevance. Not financial advice.

Originally reported at

blogs.nvidia.com

Discernion covers the story. Read the full piece at the source.

Tagsaihardwaretechbusinessllmseurope

Intelligence analysis by

Gemini 2.5 Flash

Published

Jul 21, 2026

Source

blogs.nvidia.com

Share

Topics

aihardwaretechbusinessllmseurope

Related

More from this desk

Jul 21·techcrunch.com

US threatens sanctions against Chinese AI models over IP theft

The U.S. Treasury Secretary has threatened sanctions against Chinese AI companies if intellectual property (IP) theft is found in their open-source models, marking a significant escalation in the tech rivalry.

Jul 21·deepmind.google

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google DeepMind has launched new Gemini Flash models, including 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, designed for enhanced efficiency, lower latency, and reliable performance in AI agent development.

Vector illustration of the Gemini logo.
Jul 21·theverge.com

Google launches a cheaper alternative to large AI security models like Mythos

Google has introduced Gemini 3.5 Flash Cyber, a new cost-efficient AI security model designed to quickly identify and patch vulnerabilities, positioning it as an alternative to more expensive systems like Anthropic's Mythos.

Jul 21·wired.com

Halliday’s New Smart Glasses Skip the Camera

Halliday has introduced its new smart glasses, the G2, which skip the camera and focus on listening to workplace meetings, summarizing the interesting bits, and providing AI-powered note-taking features.