← Back
Sirui-Xu

Sirui-Xu/ULTRA

[IROS 2026] ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

View on GitHub ↗
Stars
26
Forks
1
Watchers
26
Open issues
2
Contributors
1
Language
Python
License
Apache License 2.0
Default branch
main
Created Sep 25, 2026Updated Sep 27, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation

Xialin He*  Sirui Xu*  Xinyao Li  Runpei Dong  Liuyu Bian  Yu-Xiong Wang†  Liang-Yan Gui†
University of Illinois Urbana-Champaign
*Equal contribution †Equal advising
IROS 2026 · Best Paper Award on Mobile Manipulation · Finalist 🏆

🏠 Overview

ULTRA demo

ULTRA is a single multimodal controller for humanoid whole-body loco-manipulation: it tracks a motion reference when one is available, and acts from egocentric perception and a sparse goal when it is not — one policy, one set of weights, on a real Unitree G1.

⚙️ Installation

See docs/install.md: Python 3.8, pip install -r requirements.txt, IsaacGym Preview 4, and mujoco for sim2sim. Run every command below from the repository root.

📦 Data

Download Used for Extract .pt files to
InterMimic prepared OMOMO references SMPL-X retargeting input · [T, 591] InterAct/OMOMO_retarget/
ULTRA retargeted and augmented data G1 teacher/student training · [T, 630] InterAct/OMOMO_retarget_aug/

The released G1 archive contains OMOMO physics-based retargeted and augmented motions. The source archive follows InterMimic's data format.

Data layout and supported objects

Place the .pt files directly in the directories above. ULTRA includes assets for largebox, plasticbox, smallbox, and suitcase; the retargeting script selects motions for these objects.

🚀 Reproduce the pipeline

Stage Input → output Entry point
1 · Retarget Prepared OMOMO SMPL-X → G1 policy scripts/train_retarget_smplx.sh
2 · Augment SMPL-X + trained policy → G1 rollouts scripts/export_retarget_smplx.py
3 · Train teacher G1 rollouts → tracking policy scripts/train_teacher.sh
4 · Distill student G1 rollouts + included teacher → student scripts/train_student.sh
5 · Finetune student Distilled student → goal-directed policy scripts/train_finetune.sh
1 · Train the retargeting policy

After extracting the InterMimic OMOMO_retarget archive, run from the repository root:

scripts/train_retarget_smplx.sh

The script prepares supported object clips in InterAct/OMOMO_retarget_supported_080_080_080/ and trains UltraG1 with position target PD control. This data generation stage has no domain randomization, random pushes, or observation noise.

The configuration runs up to 50,000 epochs and writes checkpoints to output/retarget_smplx/g1_retarget_smplx/nn/.

2 · Export retargeted and augmented G1 motion

Use a trained retargeting checkpoint to replay the prepared references. --xyz X Y Z scales the sparse trajectories along each axis; repeat it for augmentation variants. --asset-scale selects an available object size (080_080_080 or 100_100_100).

python scripts/export_retarget_smplx.py \
  --input-dir InterAct/OMOMO_retarget \
  --checkpoint output/retarget_smplx/g1_retarget_smplx/nn/g1_retarget_smplx.pth \
  --output-dir output/retarget_export_080 \
  --asset-scale 080_080_080 \
  --xyz 1 1 1 --xyz 1.05 1 0.95

# Repeat with the larger released object assets.
python scripts/export_retarget_smplx.py \
  --input-dir InterAct/OMOMO_retarget \
  --checkpoint output/retarget_smplx/g1_retarget_smplx/nn/g1_retarget_smplx.pth \
  --output-dir output/retarget_export_100 \
  --asset-scale 100_100_100 \
  --xyz 1 1 1 --xyz 1.05 1 0.95

The exporter writes the resulting G1 motions to the chosen output directory. The released InterAct/OMOMO_retarget_aug/ archive is ready to use for teacher training.

3–4 · Train a teacher and distill the student
# Teacher — dense full-body tracking (PPO, 4096 envs). Output: output/teacher/g1_teacher/nn/
scripts/train_teacher.sh

# Student — multimodal distillation from the included teacher checkpoint.
scripts/train_student.sh
NUM_GPUS=4 scripts/train_student_multigpu.sh          # torchrun, single node

The teacher uses InterAct/OMOMO_retarget_aug/ by default. Pass --motion_file <directory> to train on another G1 motion directory.

5 · Finetune the student

Start RL finetuning from a distilled student checkpoint:

scripts/train_finetune.sh <student.pth>

The finetuning configuration uses InterAct/OMOMO_retarget_aug/ and the included tracking teacher by default.

🎮 Inference

🧠 Tracking teacher

The included checkpoint at ultra/weights/teacher_ultra_inference.pth is used for student distillation and inference. See the model card for its training data and use terms.

For the released checkpoint, we further trained on AMASS and BONES-SEED alongside OMOMO after the original work. The training code in this repository follows the original OMOMO setting.

Run teacher inference
# Put one retargeted G1 .pt clip in /absolute/path/to/one_motion_dir.
python ultra/run_teacher_inference.py --motion-dir /absolute/path/to/one_motion_dir

The command writes output/teacher_inference/rollout.pt with observed/reference states and actions from the motion. Use --checkpoint <teacher.pth> for another tracking teacher or change env.teacherPolicy in ultra/data/cfg/g1_student_vae.yaml for distillation.

🤖 Student

scripts/play_student.sh <student.pth> 1 full_track      # IsaacGym playback (also: sparse_track, object_obs points)
scripts/sim2sim_student.sh <student.pth> InterAct/OMOMO_retarget_aug full_track   # MuJoCo
scripts/export_jit.sh <student.pth> student_jit.pt      # TorchScript export

📄 License

ULTRA's own code is released under Apache-2.0 (top-level LICENSE); the InterMimic-derived simulation and learning stack is under MIT (LICENSE-InterMimic). ultra/env/tasks/base_task.py and vec_task.py are adapted from NVIDIA's IsaacGym examples and keep NVIDIA's header. Robot assets are from Unitree; motion data derives from OMOMO through InterAct.

📖 Citation

@inproceedings{he2026ultra,
  title     = {ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation},
  author    = {He, Xialin and Xu, Sirui and Li, Xinyao and Dong, Runpei and Bian, Liuyu and Wang, Yu-Xiong and Gui, Liang-Yan},
  booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year      = {2026}
}