Robots go generalist with shared foundation models

The gist
Robots powered by foundation models are learning from each other across platforms, smashing specialist performance and rewriting the rules for real-world autonomy.
What to know
- By early 2026, robots trained on massive, shared datasets from 30+ labs outperform specialized bots by 50%, unlocking cross-platform generalization.
- New models infer their own design from camera imagesdno more handcraftingdwhich turbocharges multi-robot data sharing and speeds up deployment.
- Industry pilots by Path Robotics and MagicLab Robotics show off scalable, collective robot learning, but real-world rollout is still bottlenecked by hardware limits and data generation hurdles.
Robots Learn Like the Web
Foundation models trained across diverse labs enable robots to continuously improve by sharing experiences, unlocking rapid generalization and surpassing specialized designs.
By early 2026, robotic foundation models trained on massive, diverse data sets from over 30 academic labs have demonstrated remarkable generalization, outperforming specialized models by approximately 50% across varied tasks and platforms. This collective learning approach, where robots share all experiences regardless of embodiment, enables a powerful data flywheel effect—robots continuously improve as they are deployed, creating a virtuous cycle of autonomous learning and adaptability. Sergey Levine highlights this transformative shift, comparing it to how ChatGPT learned from the internet, emphasizing that "the more variety the model sees, the better it gets."
A striking advantage of these foundation models lies in their ability to handle cross-embodiment learning with minimal explicit encoding of robot morphology. Instead of requiring handcrafted specifications, the models infer the robot type simply by analyzing camera images, simplifying integration across diverse robotic platforms. This capability not only accelerates multi-robot data sharing but also parallels breakthroughs in language models, where broad, multi-task training outperforms specialized systems, signaling a paradigm shift from perfecting individual robots to leveraging collective intelligence.
The revolution in robotics driven by multimodal foundation models extends beyond humanoid robots, unlocking unprecedented creativity and real-world deployment potential across various robot form factors. These adaptable, multimodal autonomous agents push AI-driven physical intelligence to new heights by integrating vision, language, and action, enabling robots to operate effectively in unpredictable environments. This broad generalization and performance optimization herald a new era where generalist robots surpass traditional specialized designs, fundamentally transforming autonomy and cross-platform adaptation by 2026.
From Pilots to Production
Strategic, incremental rollouts and tightly controlled failures are bridging the gap between lab breakthroughs and real-world robotics, as industry leaders deploy AI-driven machines in high-stakes environments.
Real-world deployment of robotics is rapidly advancing through strategic incremental rollouts and specialized applications that address unique operational challenges. Companies like Path Robotics are limiting initial production to 50 autonomous welding quadrupeds targeting constrained industrial environments such as shipbuilding, while MagicLab Robotics demonstrated large-scale multi-robot coordination with swarms of robot dogs and humanoids at public events, showcasing embodied AI capabilities at scale. This approach aligns with the necessity of tightly controlled, small, and safe early failures to mitigate costly disruptions, as emphasized in practical automation models that stress surfacing uncertainties before final tooling and safety certifications.
Despite impressive laboratory breakthroughs, a significant gap persists between robotics capabilities in controlled settings and robust, scalable field deployment. Experts like Ardalan Tajbakhsh of Amazon Robotics describe the transition from pilot projects to scalable solutions as 'messy and difficult,' with robots often solving only parts of use cases and requiring ongoing human intervention. This challenge is compounded by hardware constraints, including battery life and energy efficiency, which remain critical bottlenecks, especially in military and outdoor applications where safety and operational endurance are paramount.
The industrial impact of robotics is becoming tangible as robots evolve into essential enterprise assets that outperform humans in high-skill environments, with major players like NVIDIA, Siemens, and Rockwell Automation showcasing AI-driven manufacturing innovations at Hannover Messe 2026. Incremental rollout strategies are proving effective, exemplified by Lenovo’s deployment of production-scale AI delivering up to 85% faster lead times, while Bosch emphasizes the critical interplay between humans and AI to ensure safety and collaboration in industrial settings. This shift underscores the convergence of hardware and software as foundational to Industry 5.0’s promise.
Deploying physical AI systems faces fundamental hardware and safety challenges that often overshadow model intelligence, as noted by industry insiders highlighting the trade-offs between power, cost, and performance in embedded systems operating under harsh conditions. While large models like Google's Gemma 2B can technically run on embedded platforms, substantial customization is required to make them practical for robotics applications. Moreover, autonomy functions are predominantly developed in-house due to their specialized nature, whereas more generic tasks such as voice assistance leverage generalist large language models, with latency considerations critically shaping the balance between on-device processing and cloud computation.
Data-Driven Robotic Intelligence
Combining web-scale video pretraining with targeted teleoperation data and multimodal inputs, new robotics pipelines accelerate learning and physical intelligence while grappling with costly data bottlenecks.
Robotics training pipelines have innovatively combined web-scale video pretraining with targeted teleoperation data to overcome the traditional scarcity of robotics datasets. By first pretraining models on vast internet video collections using video generation objectives—predicting future frames without requiring action labels—and then fine-tuning with 10 to 20 hours of multimodal teleoperation data (vision, state, and action), systems effectively translate video predictions into executable robot actions. This two-stage approach, highlighted in analyses from April 2026, leverages high-quality filtered video data to sidestep the need for extensive action annotations, enabling scalable learning from internet-scale sources while grounding models in real-world embodied experience.
The integration of hierarchical architectures plays a pivotal role in synthesizing diverse data types for robotics learning, where internet-scale video data informs higher-level perceptual and motion primitives, and embodied teleoperation data refines lower-level control and force-sensitive interactions. As experts note, while video increasingly provides depth and free space motion cues useful for tasks like object grasping, sensing forces and physical interactions still demand embodied data enriched with tactile feedback. This layered approach acknowledges that purely predictive models common in large language architectures may fall short for embodied agents, necessitating specialized world models that act as internal simulations forecasting the consequences of robot actions.
Advances in multimodal AI, exemplified by NVIDIA’s systems, demonstrate how accepting diverse inputs—including vision, language, and action—enables robots to execute highly expressive and stable motions with remarkable efficiency, reducing the thousands of trial-and-error attempts previously needed in simulation. Jerry Pratt emphasizes that extending this multimodal integration to incorporate physical common sense through force-aware learning and potentially additional senses like sound and smell is crucial for developing robotic cognition. However, data generation remains a bottleneck; teleoperation, while rich, is costly and imperfect, underscoring the urgent need for scalable, cost-effective data collection methods and tactile sensor integration to imbue foundation models with physical intelligence that accelerates learning with less data.
Despite the promise of teleoperation data collection exemplified by the Neo robot’s strategy to amass internet-scale humanoid robot datasets, reliance on human-generated data alone is insufficient to capture the complexity and variability of robotic dynamics. The rich information embedded in position, velocity, and force data across many degrees of freedom demands a hybrid approach that merges physics-based understanding with data-driven learning. This fusion is essential to surmount the general intelligence challenge in humanoid robots, as purely empirical data cannot fully encapsulate the nuanced physical interactions required for autonomous, adaptive behavior.



