NVIDIA Cosmos 3 models arrive on Amazon SageMaker JumpStart, expanding the availability of its advanced artificial intelligence portfolio with open, frontier omnimodal world models. AWS customers can now access three specialized foundation models designed explicitly for physical AI development, empowering engineers to construct sophisticated robots, autonomous vehicles, and vision AI systems capable of perceiving, reasoning, planning, and acting within physical environments.
Deploying Frontier Physical AI Through AWS
The integration into Amazon SageMaker JumpStart provides enterprise developers and researchers with streamlined access to cutting-edge model deployment. Rather than managing complex underlying infrastructure configurations, engineering teams can deploy any of the three models with just a few clicks through the SageMaker console or programmatically utilizing the SageMaker Python SDK within their designated AWS accounts.
These models bridge the critical gap between digital intelligence and physical execution, addressing a wide array of industrial challenges that demand real-time physical reasoning, spatial perception, and high-fidelity environment simulation. Each model in the Cosmos 3 lineup targets distinct operational constraints, scaling from resource-constrained embedded systems up to massive cloud-based simulation pipelines.
The Cosmos 3 Model Portfolio Specifications
The newly available suite consists of three distinct variants tailored to specific performance envelopes and workload requirements:
- Cosmos3-Edge: Engineered explicitly for on-device robot control and real-time visual reasoning on specialized edge hardware. This 4B-parameter omni-model integrates a 2B Nemotron-based reasoner operating at robot-control resolution (640×360). It delivers real-time reasoning and generates 32 actions per inference at 15 Hz on NVIDIA Jetson Thor, while supporting 256p and 480p video streams at 12 to 30 FPS.
- Cosmos3-Nano: A compact 16B-parameter omnimodal model optimized for physics-aware world generation and deep physical reasoning. It processes multi-modal inputs combining text, images, video, audio, and action trajectories to produce synchronized outputs, allowing agents to utilize prior knowledge, common sense, and physical intuition. It supports advanced chain-of-thought reasoning over text, images, and video at resolutions up to 720p.
- Cosmos3-Super: The premier model in the family, providing highest-fidelity world generation and simulation at 64B parameters. Utilizing a unified Mixture-of-Transformers architecture, it jointly processes and generates language, images, video, audio, and action sequences across multiple aspect ratios up to 720p resolution, making it ideal for large-scale simulation, synthetic data generation, and policy learning workflows.
Practitioners seeking to integrate these capabilities can reference the comprehensive documentation available through the Amazon SageMaker JumpStart model catalog to configure their deployment pipelines effectively.
Source: Original Article





