The model functions as a closed-loop system where robots generate candidate actions based on language instructions and visual history, predict the resulting states, and evaluate whether those outcomes achieve specific goals. During execution, Motus2 employs Best-of-N planning to compare potential actions in real-time, executing the option with the highest predicted success score. This architecture effectively bridges the gap between pre-trained human manipulation knowledge and physical robotic control.
Performance benchmarks highlight the system's efficacy. In real-robot trials involving tasks such as screwing in a light bulb and multi-finger manipulation, Motus2 achieved an average success rate of 84%. Researchers observed that incorporating a lightweight tactile expert, which processes contact feedback to refine movements, improved performance on tasks like tearing paper by 12.5 percentage points. The model is trained on a massive foundation of 130,000 hours of egocentric human video data, subsequently refined with over 100 hours of robot-specific trajectories.
ShengShu Technology positions Motus2 as a key implementation of its L3 roadmap, representing a shift toward robots that can actively learn from their own successes and failures. While the current iteration focuses on specific dexterous objectives, the company intends to use this framework to advance toward autonomous agents capable of long-term reasoning in varied environments. The research, architecture, and demonstration videos are currently available for public review.

Comments (0)
No comments yet. Be the first!