← Back
shengshu-ai

shengshu-ai/Motus2

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

View on GitHub ↗
Stars
397
Forks
2
Watchers
397
Open issues
0
Contributors
1
Language
—
License
—
Default branch
main
Created Sep 11, 2026Updated Sep 11, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

Project Page arXiv Hugging Face

Overview

General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world–action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human–robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.

Motus2 overview

Updates

  • [2026-09] We plan to release the Motus2 code and checkpoints progressively throughout September 2026.

Open-Source Roadmap

  • Stage 1 pretraining checkpoints
  • Stage 2 pretraining checkpoints
  • Video pretraining code (Stage 1)
  • Video–Action (Value) training code (Stage 2 pretraining, mid-training, and post-training)
  • MBRL code
  • Memory and variable-length training infrastructure
  • Tactile code

Links

  • Project page
  • Paper
  • Models

Citation

If you find Motus2 useful in your research, please cite:

@misc{bi2026motus2,
  title         = {Motus2: A Self-Evolving General World Model for Dexterous Manipulation},
  author        = {Hongzhe Bi and Zihao Zhou and Yihang Tang and Jingrui Pang and
                   Shuhe Huang and Haitian Liu and Runqing Wang and Shuai Huang and
                   Yichen Wang and Yiming Cheng and Ruowen Zhao and Zhenghua Li and
                   Hengkai Tan and Xiaolong Liu and Jinhui Wan and Jiabao Liu and
                   Min Zhao and Fan Bao and Jun Zhu},
  year          = {2026},
  eprint        = {2608.30237},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2608.30237}
}