Xialin He*
Sirui Xu*
Xinyao Li
Runpei Dong
Liuyu Bian
Yu-Xiong Wang†
Liang-Yan Gui†
University of Illinois Urbana-Champaign
*Equal contribution †Equal advising
IROS 2026 · Best Paper Award on Mobile Manipulation · Finalist 🏆
ULTRA is a single multimodal controller for humanoid whole-body loco-manipulation: it tracks a motion reference when one is available, and acts from egocentric perception and a sparse goal when it is not — one policy, one set of weights, on a real Unitree G1.
See docs/install.md: Python 3.8, pip install -r requirements.txt, IsaacGym Preview 4, and
mujoco for sim2sim. Run every command below from the repository root.
| Download | Used for | Extract .pt files to |
|---|---|---|
| InterMimic prepared OMOMO references | SMPL-X retargeting input · [T, 591] |
InterAct/OMOMO_retarget/ |
| ULTRA retargeted and augmented data | G1 teacher/student training · [T, 630] |
InterAct/OMOMO_retarget_aug/ |
The released G1 archive contains OMOMO physics-based retargeted and augmented motions. The source archive follows InterMimic's data format.
Data layout and supported objects
Place the .pt files directly in the directories above. ULTRA includes assets for largebox, plasticbox, smallbox, and suitcase; the retargeting script selects motions for these objects.
| Stage | Input → output | Entry point |
|---|---|---|
| 1 · Retarget | Prepared OMOMO SMPL-X → G1 policy | scripts/train_retarget_smplx.sh |
| 2 · Augment | SMPL-X + trained policy → G1 rollouts | scripts/export_retarget_smplx.py |
| 3 · Train teacher | G1 rollouts → tracking policy | scripts/train_teacher.sh |
| 4 · Distill student | G1 rollouts + included teacher → student | scripts/train_student.sh |
| 5 · Finetune student | Distilled student → goal-directed policy | scripts/train_finetune.sh |
1 · Train the retargeting policy
After extracting the InterMimic OMOMO_retarget archive, run from the repository root:
scripts/train_retarget_smplx.shThe script prepares supported object clips in InterAct/OMOMO_retarget_supported_080_080_080/ and trains UltraG1 with position target PD control. This data generation stage has no domain randomization, random pushes, or observation noise.
The configuration runs up to 50,000 epochs and writes checkpoints to output/retarget_smplx/g1_retarget_smplx/nn/.
2 · Export retargeted and augmented G1 motion
Use a trained retargeting checkpoint to replay the prepared references. --xyz X Y Z scales the sparse trajectories along each axis; repeat it for augmentation variants. --asset-scale selects an available object size (080_080_080 or 100_100_100).
python scripts/export_retarget_smplx.py \
--input-dir InterAct/OMOMO_retarget \
--checkpoint output/retarget_smplx/g1_retarget_smplx/nn/g1_retarget_smplx.pth \
--output-dir output/retarget_export_080 \
--asset-scale 080_080_080 \
--xyz 1 1 1 --xyz 1.05 1 0.95
# Repeat with the larger released object assets.
python scripts/export_retarget_smplx.py \
--input-dir InterAct/OMOMO_retarget \
--checkpoint output/retarget_smplx/g1_retarget_smplx/nn/g1_retarget_smplx.pth \
--output-dir output/retarget_export_100 \
--asset-scale 100_100_100 \
--xyz 1 1 1 --xyz 1.05 1 0.95The exporter writes the resulting G1 motions to the chosen output directory. The released InterAct/OMOMO_retarget_aug/ archive is ready to use for teacher training.
3–4 · Train a teacher and distill the student
# Teacher — dense full-body tracking (PPO, 4096 envs). Output: output/teacher/g1_teacher/nn/
scripts/train_teacher.sh
# Student — multimodal distillation from the included teacher checkpoint.
scripts/train_student.sh
NUM_GPUS=4 scripts/train_student_multigpu.sh # torchrun, single nodeThe teacher uses InterAct/OMOMO_retarget_aug/ by default. Pass --motion_file <directory> to train on another G1 motion directory.
5 · Finetune the student
Start RL finetuning from a distilled student checkpoint:
scripts/train_finetune.sh <student.pth>The finetuning configuration uses InterAct/OMOMO_retarget_aug/ and the included tracking teacher by default.
The included checkpoint at ultra/weights/teacher_ultra_inference.pth is used for student distillation and inference. See the model card for its training data and use terms.
For the released checkpoint, we further trained on AMASS and BONES-SEED alongside OMOMO after the original work. The training code in this repository follows the original OMOMO setting.
Run teacher inference
# Put one retargeted G1 .pt clip in /absolute/path/to/one_motion_dir.
python ultra/run_teacher_inference.py --motion-dir /absolute/path/to/one_motion_dirThe command writes output/teacher_inference/rollout.pt with observed/reference states and actions from the motion. Use --checkpoint <teacher.pth> for another tracking teacher or change env.teacherPolicy in ultra/data/cfg/g1_student_vae.yaml for distillation.
scripts/play_student.sh <student.pth> 1 full_track # IsaacGym playback (also: sparse_track, object_obs points)
scripts/sim2sim_student.sh <student.pth> InterAct/OMOMO_retarget_aug full_track # MuJoCo
scripts/export_jit.sh <student.pth> student_jit.pt # TorchScript exportULTRA's own code is released under Apache-2.0 (top-level LICENSE); the InterMimic-derived simulation and learning
stack is under MIT (LICENSE-InterMimic). ultra/env/tasks/base_task.py and vec_task.py are adapted from
NVIDIA's IsaacGym examples and keep NVIDIA's header. Robot assets are from Unitree; motion data derives from OMOMO
through InterAct.
@inproceedings{he2026ultra,
title = {ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation},
author = {He, Xialin and Xu, Sirui and Li, Xinyao and Dong, Runpei and Bian, Liuyu and Wang, Yu-Xiong and Gui, Liang-Yan},
booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
year = {2026}
}