Nvidia cosmos 3 sets new bar for open physical AI

Drip

The gist

NVIDIA's Cosmos 3 open foundation models redefine physical AI, unifying vision, language, and action prediction for real-time reasoning across robotics, autonomous vehicles, and beyond.

What to know

Unified AI for Physical World

Cosmos 3’s mixture-of-transformers architecture fuses multimodal perception and action prediction, enabling zero-shot generalization and real-time control on everything from data centers to compact edge devices.

NVIDIA's Cosmos 3 model family represents a groundbreaking leap in physical AI by unifying scene understanding, synthetic data generation, and action prediction within a single open foundation model. Employing a sophisticated mixture-of-transformers architecture, Cosmos 3 efficiently routes multimodal inputs—ranging from text and images to video, ambient sound, and physical actions—through specialized sub-networks, enabling robust real-time reasoning and control. This design not only supports complex world dynamics modeling but also facilitates zero-shot generalization to new tasks and environments, a capability underscored by the introduction of world action models that convert pretrained video backbones directly into robot policies, as exemplified by the compact 4-billion-parameter Cosmos 3 Edge optimized for deployment on NVIDIA Jetson and RTX GPUs.

A key technical innovation in the Cosmos 3 family is its remarkable efficiency and scalability, enabling deployment across a wide spectrum of hardware—from high-throughput NVIDIA DGX systems to real-time edge devices like Jetson Orin and Jetson Thor. Notably, NVIDIA compressed the previously cluster-dependent Cosmos-Dreams autonomous vehicle simulator, which once required 64 GB300 GPUs, to run photorealistic driving scenarios frame-by-frame on a single RTX PRO 6000 workstation GPU. This efficiency extends to the Cosmos 3 Edge model, which operates at 15 Hz on compact modules, demonstrating NVIDIA's commitment to making physical AI pipelines accessible and practical for diverse real-world applications.

NVIDIA’s open licensing strategy, implemented through the Linux Foundation’s OpenMDW 1.1 license, is pivotal in fostering transparency and collaborative innovation in physical AI. By releasing not only the Cosmos 3 model weights but also inference code and post-training recipes, NVIDIA empowers developers to adapt and specialize models to their unique robots, vehicles, sensors, and operational environments—an approach NVIDIA frames as a practical necessity rather than a mere licensing preference. This openness is complemented by integration with Omniverse libraries and OpenUSD frameworks, which streamline the creation of simulation-ready environments and efficient management of 3D assets and sensor configurations, thereby accelerating development cycles across robotics, autonomous vehicles, and vision AI domains.

Benchmark evaluations consistently position Cosmos 3 at the forefront of open physical AI models, achieving top rankings across diverse tasks including text-to-image and image-to-video generation, world generation, robot policy learning, and vision understanding. For instance, Cosmos 3 leads Artificial Analysis benchmarks for open-weights generative tasks, tops PAI-Bench for world generation, dominates Physics-IQ’s image-to-video category, and secures first place on RoboLab for robot policy, with the Cosmos 3 Super variant recognized as the highest-ranked open model on VANTAGE-Bench for vision analytics. These results underscore the model family’s state-of-the-art capabilities and its broad applicability across physical AI challenges.

Sources

Open Models, Local Power

NVIDIA’s open ecosystem and regional post-training strategy let enterprises and countries fully customize and govern AI models, ensuring sovereignty and safety without the need for massive compute resources.

NVIDIA’s ecosystem coalitions, notably Nemotron and Cosmos, exemplify its strategic commitment to open foundation models as a long-term initiative akin to CUDA’s transformative impact. By inviting partners into the model-building process early, NVIDIA fosters collaborative innovation that ensures models are tailored to diverse industry needs and regional contexts. This approach not only accelerates adoption but also empowers enterprises and countries to own, inspect, and customize AI models with their own data, reinforcing safety, cybersecurity, and sovereignty as emphasized by CEO Jensen Huang.

Recognizing the unique toolsets and requirements of different enterprises and regions, NVIDIA advances regional adaptation through post-training on localized data, enabling AI models to integrate seamlessly with up to thousands of local tools. This nuanced customization is supported by a tiered model offering—Nano, Super, and Ultra—that accommodates varying hardware capabilities, allowing developers in resource-constrained environments to iterate efficiently before scaling up. Such flexibility ensures that even regions lacking extensive compute resources can bootstrap sovereign AI capabilities without the prohibitive cost of recreating foundational datasets.

NVIDIA’s open model ecosystem bridges the symbolic and physical AI domains, spanning from language and code-based reasoning to real-world applications in robotics and autonomous vehicles. By releasing not only model weights but also the underlying data openly, NVIDIA accelerates cross-domain innovation and adoption, enabling industries to leverage these foundational tools for both digital and physical AI challenges. This comprehensive openness reflects lessons learned from shipping in the open and underscores NVIDIA’s vision of fostering a diverse, adaptable AI ecosystem.

Sources

Industrial Impact in Action

Cosmos 3 is already transforming industries—powering high-speed surgical simulators, robust robotics training, and next-gen autonomous driving with open, scalable foundation models.

NVIDIA's Cosmos 3 models are driving transformative real-world industrial deployments across multiple sectors, from vision AI agents accelerated by Nvidia Metropolis to advanced surgical simulators. Notably, the Cosmos-H-Dreams surgical simulator operates live at 160 frames per second on a single RTX PRO 6000 GPU, enabling real-time interaction with robotic arms on CMR Surgical’s Versius platform, supported by nearly 500 hours of anonymized surgical data. This generative simulation approach complements traditional physics engines, positioning NVIDIA’s Isaac for Healthcare stack as a cutting-edge research tool that enhances surgical robotics development while respecting regulatory boundaries.

In robotics, NVIDIA Cosmos 3 underpins scalable and safe training environments that closely mirror real-world conditions, enabling early-stage companies to accelerate model evaluation with unprecedented control over physical parameters like lighting, friction, and object geometry. As one client emphasizes, the strong alignment between simulation and reality instills confidence, allowing for faster, cost-effective iteration cycles that overcome the traditional bottlenecks of physical testing. This digital infrastructure is rapidly gaining industrial traction by facilitating robust robot learning and pragmatic downstream applications.

NVIDIA’s open-source Alpamayo 2 Super, a 34-billion-parameter Vision-Language-Action foundation model built on Cosmos 3, is revolutionizing autonomous driving by delivering deep environmental understanding and predictive capabilities beyond conventional perception systems. Freely available under the OpenMDW-1.1 license on Hugging Face, Alpamayo 2 Super compresses data annotation timelines from months to days and integrates seamlessly into NVIDIA’s comprehensive autonomous driving stack—including Hyperion hardware, Halos OS software, and simulation tools like Omniverse—enabling automakers to customize and deploy Level 4 autonomous systems with continuous closed-loop training and reinforcement learning.

TIER IV’s deployment of NVIDIA Cosmos models within its Co-MLOps platform exemplifies industrial adoption of physical AI to tackle the long-tail problem in autonomous driving through automated data search, edge-case generation, and multi-modal data augmentation. Leveraging a Collaborative Multi-stage Ensemble-based Teacher Model (CoMET) that integrates 12 large-scale models, the platform efficiently autolabels diverse sensor inputs—including LiDAR and multi-FoV cameras—across challenging scenarios such as tunnels, construction zones, and adverse weather, enabling safe autonomous driving demonstrations across 39 prefectures in Japan and showcasing the scalability and robustness of NVIDIA’s physical AI ecosystem.

Sources

Toward True Spatial Intelligence

Cosmos 3’s world action models and synthetic data pipelines drive AI systems that can reason about, predict, and interact with complex physical environments—bridging the gap between digital intelligence and real-world autonomy.

NVIDIA's Cosmos 3 open foundation model family exemplifies the cutting edge of multimodal physical AI by integrating vision reasoning, language understanding, video world modeling, and action prediction into a unified Mixture-of-Transformers architecture. This omni-model approach enables advanced spatial intelligence, allowing AI systems to not only interpret but also predict and simulate how physical environments evolve, which is crucial for robotics and autonomous vehicles. As Fei-Fei Li emphasizes, this new frontier of AI focuses on three-dimensional reasoning that bridges physical and virtual spaces, empowering robots to understand and act effectively within complex real-world settings.

The development of World Action Models (WAMs) built on Cosmos 3’s video world models marks a significant leap beyond traditional vision-language-action frameworks by embedding physical dynamics and physics-based reasoning into AI policies. This enables zero-shot transfer learning, allowing robots to generalize across new tasks, embodiments, and environments with dramatically less task-specific data, as these models inherently understand how objects move and interact. NVIDIA’s technical blog highlights that such physics-aware models reduce reliance on extensive retraining and accelerate deployment across diverse robotic platforms.

Complementing these AI advances, NVIDIA’s ecosystem leverages Omniverse and OpenUSD to create scalable, simulation-ready environments that facilitate a 'real to sim to real' pipeline. This approach, championed by Yun Zhu Li, replaces costly and limited real-world data collection with synthetic data generation and rigorous policy testing in digital twins, effectively addressing rare events and long-tail scenarios that are otherwise difficult to capture. By streamlining data reuse and environment configuration, this infrastructure accelerates physical AI development and deployment across robotics, autonomous vehicles, and vision AI systems.

The versatility of the Cosmos 3 model family—from the high-fidelity Cosmos 3 Super (64B) to the edge-optimized Cosmos 3 Edge (4B)—enables deployment across a broad spectrum of hardware, including NVIDIA RTX GPUs, DGX systems, and Jetson platforms like Jetson Thor. This hardware adaptability ensures that specialized robot policies and vision reasoning can operate efficiently at the edge, meeting the real-time demands of physical AI applications in autonomous vehicles and robotics. This scalability is critical for translating multimodal AI capabilities from research prototypes to practical, industrial-grade systems.

Sources

Part of these trends

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.