Robots get a feel for the future: tactile AI surges ahead

The gist
Robots are finally learning to 'feel' their world, with next-gen tactile AI turning science fiction dexterity into real-world breakthrough.
What to know
- BeingBeyond’s Being-H0.8 model lets robots combine touch and vision for far more natural, nuanced physical interactions.
- Daimon Robotics just landed several hundred million yuan from Ant Group and others to accelerate its tactile-powered ‘Physical Interaction Brain’ for dexterous robot control.
- New models like N_0-VTLA, trained on massive tactile datasets, are hitting up to 95% success rates in complex real-world tasks—signaling a seismic shift from vision-only to touch-anchored robotics.
Touch: The Next AI Frontier
Robots are gaining the power to 'understand by touch,' integrating tactile intelligence with vision to interact with the physical world as intuitively as humans.
BeingBeyond's launch of the Being-H0.8 model marks a significant leap in tactile world-action systems, enabling robots to integrate touch and vision for a richer understanding of their surroundings. This new generation model emphasizes tactile intelligence as a core component of embodied AI, allowing machines not just to see but to 'understand by touch,' which enhances their ability to interact with complex environments more naturally and effectively.
China’s Tactile AI Bet
Massive funding from Ant Group and others is fueling a race to build tactile-powered robot brains, signaling a strategic shift toward physical interaction as the key to real-world dexterity.
Daimon Robotics has secured a strategic financing round worth several hundred million yuan led by Ant Group, with participation from prominent investors such as China Merchants Capital, Lenovo Capital, China Mobile, and China Telecom. This substantial funding injection is aimed at accelerating the development of integrated vision-tactile-language-action (VTLA) systems and the world’s first 'Physical Interaction Brain,' which combines physical cognition, deductive decision-making, and real-time control to enable dexterous robot operation. The company’s commercialization strategy is underpinned by a comprehensive tactile intelligence infrastructure, including mass-producible tactile sensors, extensive physical interaction data, and deployment systems, positioning Daimon at the forefront of tactile embodied AI innovation and market readiness.
Ant Group’s strategic investment in Daimon Robotics signals a forward-looking bet on tactile sensing as a critical frontier in embodied AI, moving beyond traditional vision-language models to enhance robot reliability in contact-intensive tasks. This shift demands a fundamental redesign of data architectures, model frameworks, and control methodologies, as evidenced by Daimon’s launch of the Daimon-TWM tactile-anchored world model that integrates tactile feedback across physical state understanding, action prediction, and real-time control. By backing this tactile-anchored ecosystem, Ant Group aims to catalyze a transformation in robotics from merely 'walking and seeing' to truly 'working' with sophisticated physical interaction capabilities, despite the inherent uncertainties in this emerging field.
Scaling Up Robot Intelligence
Large-scale tactile datasets and next-gen world-action models are smashing previous limits, driving robots to human-like mastery in complex, contact-rich tasks.
The N_0-VTLA model marks a significant milestone as the first vision-tactile-language-action system pretrained at scale on tactile data, leveraging the NeoData dataset to excel in contact-rich manipulation tasks. Its innovative training approach, including the ALTER offline reinforcement learning method, enables remarkable success rates of 75-95% on complex real-robot tasks, outperforming strong baselines by a wide margin—winning all nine NeoReal tasks and achieving 63.8% mean success on a twenty-task simulation suite compared to 44.0% for competitors. This demonstrates the critical role of large-scale visuo-tactile pretraining combined with staged tactile-pathway integration in advancing embodied AI capabilities.
DYNA-2, developed by Dyna Robotics, introduces a paradigm-shifting World-Action Model (WAM) architecture that jointly predicts future video frames and motor commands, effectively endowing robots with spatial reasoning and contact physics understanding absent in traditional Vision-Language-Action (VLA) models. Trained on over one million hours of egocentric human video—equivalent to 170 years of continuous experience—DYNA-2 establishes a human-to-robot scaling law, showing smooth performance improvements across four orders of magnitude in training data without plateauing. This approach dramatically reduces reliance on costly teleoperated robot data, enabling rapid adaptation to new tasks with as little as 13 minutes of fine-tuning, and achieves up to an 87% pass rate in real-world deployments, significantly outperforming its VLA-based predecessor.
Scaling embodied AI models hinges on overcoming fundamental data and embodiment challenges, as robotics is intrinsically a data problem requiring vast, diverse datasets spanning simulation, physical demonstrations, and real-world robot operation. The 'cold start' problem in data acquisition is being mitigated by leveraging rapidly growing internet-scale and egocentric datasets, particularly fueled by the surge in robot deployments in China and elsewhere. Moreover, recent shifts in robot control philosophies—such as open access to motor and control systems exemplified by Franka robots—have accelerated progress by enabling tighter integration of models with hardware, while emphasizing that ultimate model validation occurs only through physical robot deployment.
A critical bottleneck in embodied AI development lies in defining and aligning the observation and action spaces—the sensory inputs like vision and tactile feedback and the physical capabilities to interact with environments. Many companies have historically guarded their data, limiting progress, but addressing these embodiment constraints is essential for effective model training and deployment. This recognition complements foundational research efforts like N_0-VTLA and DYNA-2, which explicitly incorporate tactile information and physical interaction modeling, underscoring that successful embodied AI requires not just data scale but also careful consideration of the robot’s sensory and motor embodiment.
Humans in the Tactile Loop
Human demonstration, prosthetic data, and physics-driven models are fusing to teach robots nuanced manipulation—accelerating deployment and bridging the gap between human and machine skills.
Reimagine Robotics exemplifies a human-in-the-loop paradigm that enhances embodied AI by enabling workers to train robots through direct demonstration and real-time correction, a process CEO Jonathan Scholz describes as 'monkey-see, monkey-do.' This approach not only accelerates deployment—cutting prototyping and testing time from a day to about 10 minutes—but also fosters collaborative workflows where robots amplify rather than replace human labor, as evidenced in diverse industrial tasks like 3D printer tending and electronics disassembly that integrate vision, tactile feedback, and human guidance.
The fusion of human-generated prosthetic data with compliant robotic hardware, as seen in ABB Robotics’ collaboration with PSYONIC, advances robotic dexterity by capturing nuanced touch and motion signals from the PSYONIC Ability Hand to train cobots for delicate, variable tasks. Marc Segura of ABB highlights this as a critical step toward closing the gap between human and robot manipulation skills, enabling applications across automotive, aerospace, packaging, and life sciences where handling fragile or irregular objects demands multimodal integration of myoelectric control, tactile sensing, and AI-driven learning.
World Action Models (WAMs), championed by NVIDIA’s Cosmos 3 foundation model, represent a transformative innovation by embedding physics-based world dynamics into embodied AI, allowing robots to generalize manipulation skills across new tasks and environments with less task-specific data. Unlike traditional vision-language-action models, WAMs leverage diverse interaction data to learn object dynamics, reducing data collection costs and enabling more flexible robot learning that transcends semantic perception to incorporate a deeper understanding of physical interactions.
Emerging embodied AI research and commercialization efforts, such as ITMO’s interdisciplinary Laboratory of Embodied Intelligence and Xiaomi’s open-source Robotics-1 foundation model, emphasize scalable, multimodal integration of vision, language, and action to enhance robot autonomy and adaptability without reliance on specialized sensors. Complementing these data-driven approaches, Acorn Robot’s 'instinct-driven' Natus AGE-0 model pioneers a zero-data tactile perception strategy that enables real-time, reflexive responses to physical stimuli, achieving millisecond-level adaptation and commercial success in flexible manufacturing—highlighting a broader ecosystem shift toward integrated, data-efficient embodied AI solutions that extend beyond robotics research into practical industrial applications.



