📊 Key Data
  • 84% success rate: Motus2's average success rate across benchmark tasks after fine-tuning with robot-specific data.
  • 130,000 hours of human video: Dataset used to train Motus2's foundational knowledge of object manipulation.
  • $293 million in funding: Secured by ShengShu Technology from investors like Alibaba Cloud and Baidu Ventures.
🎯 Expert Consensus

Experts would likely conclude that Motus2 represents a significant advancement in robotic dexterity, bridging the gap between digital intelligence and physical competence through self-evolving AI models and hybrid human-robot training data.

about 6 hours ago
Beyond Mimicry: How Motus2 Gives Robots a Sense of Touch and Purpose

Beyond Mimicry: How Motus2 Gives Robots a Sense of Touch and Purpose

SINGAPORE – September 14, 2026 – At the recent Inclusion Conference on the Bund, a robot screwed in a lightbulb. While that might sound like a familiar trope from science fiction, the technology driving the action, unveiled by ShengShu Technology, represents a significant departure from mere automation. The model, named Motus2, is not just executing a pre-programmed sequence; it is demonstrating a form of robotic intuition, a learned dexterity that could fundamentally reshape our physical world.

The demonstration, one of several showing the AI successfully tearing paper, opening a can, and even finding a hidden object, is the public debut of what the company calls a "self-evolving general world model." It's a dense phrase, but it boils down to a simple, powerful idea: an AI that can learn the physics and feel of the world, predict the consequences of its actions, and improve itself over time. For industries from manufacturing to logistics, and for anyone tracking the real-world application of AI, this move from digital intelligence to physical competence is a watershed moment.

The Architecture of Dexterity

For decades, the challenge with robotics hasn't just been movement, but manipulation. Picking up a solid, uniform box is one thing; delicately handling a flimsy paper cup or applying the right torque to a lightbulb is another. These tasks require a nuanced understanding of objects, spatial relationships, and force. Motus2 tackles this by unifying three critical functions into a single, cohesive AI brain.

First, there's a policy interface that generates actions, essentially asking, "Based on the goal and what I see, what should I do next?" Second, a simulator interface predicts the future, asking, "If I perform this action, what will happen?" And third, an evaluator interface judges the outcome, asking, "Does this predicted future get me closer to my goal?"

This isn't just a linear process; it's a closed loop. During execution, Motus2 rapidly imagines several possible actions, simulates their outcomes, and chooses the one its evaluator scores highest. After acting, it takes in new visual data from the real world and repeats the cycle. This "Best-of-N" planning allows the robot to navigate the uncertainties of physical interaction in real time. Crucially, this loop also drives the model's "self-evolution." By analyzing the value scores of its imagined actions, the system refines its own policy through model-based reinforcement learning. It learns from its successes, but just as importantly, it learns from its imagined failures without the costly and time-consuming need for constant physical trial and error. This internal feedback loop is what separates Motus2 from models that simply mimic, allowing it to genuinely improve its skills on the fly.

Learning from Our Hands

The secret to Motus2's sophisticated dexterity isn't just clever architecture; it's the data it was raised on. The model's foundational knowledge comes from an enormous dataset of approximately 130,000 hours of egocentric human video—recordings from a first-person perspective showing hands interacting with the world. By analyzing this vast library of human experience, the AI first learns the broad patterns of how objects move, bend, and change when manipulated.

However, learning from human video alone is not enough. A model trained only on this data achieved a respectable, but not revolutionary, 51% success rate on key tasks. The critical next step was adaptation. The team at ShengShu Technology mid-trained the model using more than 100 hours of actual robot data. This process bridges the gap between human anatomy and a robotic manipulator, adapting the broad priors learned from human hands to the specific observation and control spaces of a machine.

The results of this hybrid approach are striking. After being fine-tuned on the robot-specific data, Motus2's average success rate across five benchmark tasks—including placing a ball, multi-finger manipulation, and screwing in that lightbulb—jumped from 51% to 84%. This highlights a powerful new paradigm for training physical AI: provide a massive foundation of human experience, then refine it with domain-specific robotic data. It's a strategy that proves that to build a machine that can act in our world, it first needs to watch and learn from us.

A Touch of Reality

Vision, even when combined with predictive modeling, has its limits. A camera can't tell you the precise moment a paper cup will slip from a gripper or the exact force needed to tear a sheet of paper. To solve this, Motus2 incorporates a lightweight tactile expert, a module that adds a sense of touch to the robot's perception.

This isn't a completely separate system. The tactile module cleverly reuses intermediate computations from the main visual model, allowing it to process high-frequency contact feedback and refine an action just milliseconds before it's executed. It’s the robotic equivalent of the subtle adjustments we make with our fingers when handling a delicate object. The impact is significant. In tasks like pulling out a nested paper cup and tearing paper with two hands—jobs where contact and friction are everything—adding the tactile expert boosted the average success rate from 60% to 72.5%. This integration of vision and touch moves the technology from simply seeing the world to truly interacting with it.

The Road to Autonomous Agents

While the technical feats are impressive, it's crucial to place them within a strategic context. ShengShu Technology, founded in 2023, has rapidly positioned itself as a major contender in the race for general AI, securing an impressive $293 million in funding from heavyweights like Alibaba Cloud, Baidu Ventures, and Ant Group. Motus2 is not an isolated project but a key milestone in the company's ambitious five-level roadmap for creating "Autonomous World Agents."

Motus2 represents a concrete implementation of L3, "Acting in the World." By enabling a robot to physically manipulate its environment to achieve goals, it lays the groundwork for L4, "Autonomous World Agents," which would involve more complex task decomposition and long-term planning. This places the young company in direct competition with the robotics divisions of global giants like Google DeepMind and the advanced humanoid projects at Boston Dynamics.

The company is transparent about the road ahead. Reliable performance on longer, more complex tasks and autonomous learning in completely open, unstructured environments remain significant challenges. Yet, by making its research paper, model architecture, and demonstrations publicly available, ShengShu Technology is not just showcasing its own progress but is also providing a powerful catalyst for the entire field. The journey from a robot that can place a ball in a bowl to one that can assemble a product or assist in a hospital is long, but the dexterity and learning capabilities embodied in Motus2 mark a clear and decisive step forward.

Topics & Related

Event:
Product Launch
Theme:
Artificial Intelligence
Sector:
AI & Machine Learning
Robotics & Automation
Product:
AI & Software Platforms

📝 This article is still being updated

Are you a relevant expert who could contribute your opinion or insights to this article? We'd love to hear from you. We will give you full credit for your contribution.

Contribute Your Expertise →
UAID: 49969