Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action
NVIDIA says Cosmos 3 is an open omni-model for physical AI that unifies generation, reasoning, and action. It is now on Hugging Face with model cards, licensing, and supporting tools.
Intelligence analysis by GPT-5.4 Mini
NVIDIA is releasing Cosmos 3 as a single model for physical AI tasks instead of separate systems for generation, understanding, and policy. The launch includes two model sizes, Diffusers integration, post-training scripts, and synthetic data assets for robotics and driving use cases.
NVIDIA made a new computer brain called Cosmos 3. It is built to help machines understand the real world, make videos of the real world, and choose actions in it.
Before, different jobs needed different tools. Cosmos 3 tries to do them together, like one Swiss Army knife instead of a whole toolbox.
The company says this could help with robots, self-driving cars, and training data for tricky real-world scenes, like a warehouse or a road with unexpected trouble.
Analysis
What shipped
NVIDIA says Cosmos 3 is a new world foundation model for physical AI, and that it is now available on Hugging Face. The release bundles Cosmos 3 Super and Cosmos 3 Nano, plus model cards, licensing information, Diffusers integration for generation pipelines, post-training scripts on GitHub, and open synthetic data generation datasets.
What changed
The company frames Cosmos 3 as a shift from separate models for separate jobs. Earlier Cosmos releases split capabilities across products like world generation, controlled generation, scene understanding, and policy generation. Cosmos 3 instead uses a single omni-model built on a Mixture-of-Transformers architecture, which NVIDIA says can handle text, image, video, audio, and action in one forward pass.
How it works
According to the article, each modality is encoded separately and projected into a shared representation space. The input sequence is then split into two parts: an autoregressive stream for reasoning and understanding, and a diffusion stream for generation. NVIDIA says this lets one model behave like a vision-language model, a video generator, a dynamics model, or a robot policy without changing the architecture.
Model sizes and use cases
Cosmos 3 Nano is described as an 8B reasoner plus 8B generator aimed at efficient inference on workstation-class hardware such as an RTX PRO 6000 GPU. Cosmos 3 Super is a 32B reasoner plus 32B generator aimed at large-scale synthetic data generation and research on Hopper and Blackwell GPUs.
The article positions the system for robotics, autonomous driving, warehouse safety, and other physical-world tasks where motion, causality, and spatial relationships matter. NVIDIA also includes prompt guidance for video and action generation, emphasizing detailed narrative prompts for video and concise spatial instructions for action tasks.
Key points
- NVIDIA says Cosmos 3 is an open omni-model for physical AI reasoning and action.
- The release includes Cosmos 3 Super and Cosmos 3 Nano on Hugging Face.
- NVIDIA says the model combines world generation, physical reasoning, and action generation in one system.
- The article says the architecture uses a Mixture-of-Transformers design with autoregressive and diffusion subsequences.
- Supporting materials include Diffusers integration, post-training scripts, and synthetic data datasets.



