TL;DR
NVIDIA Cosmos is an open platform providing world models, datasets, and tools for Physical AI development. Cosmos 3 is an omnimodal world model family that integrates language, image, video, audio, and action sequences.
Key features
Omnimodal world model: Mixture-of-Transformers architecture integrates and generates language, image, video, audio, and action sequences
Reasoner surface: Performs world understanding, physical reasoning, task planning, and action prediction from text and vision inputs
Generator surface: Generates images, videos, synchronized sound, and action-conditioned rollouts from text, vision, sound, and action inputs
Action modeling: Predicts policy actions, inverse dynamics, and forward dynamics for robotics, camera movement, autonomous driving, etc.
Research and production path: Develop with Diffusers/Transformers, then serve with OpenAI-compatible vLLM-Omni/vLLM
Post-training support: Fine-tuning and customization capabilities
When to use it
When developing Physical AI systems such as robots, autonomous vehicles, and smart infrastructure that require world models
When handling omnimodal tasks like video understanding and generation, physical reasoning, and action prediction in an integrated manner
When needing world simulation for synthetic data generation, policy learning, and robot training