| Real-time inference results (single-image input) | Real-time inference results (multi-image input) |
|---|---|
![]() |
![]() |
The examples show real-time inference results with a single input image (left) and four input images (right).
- Python 3.10
- CUDA 12.6
1. Create conda environment:
conda env create -f environment.yml
conda activate inspatio_world_test2. Install Depth Anything 3 for the depth step:
python -m pip install --no-deps depth-anything-3==0.1.1environment.yml includes the packages DA3 uses on this inference path. The separate command avoids installing its unrelated app and benchmark dependencies.
On Hopper GPUs, we recommend installing FlashAttention-3 (FA3) for faster attention.
Download the following model checkpoints into the checkpoints/ directory:
| Model | Purpose | Source |
|---|---|---|
| InSpatio-World-1.5 | Default v2v checkpoint — 1.3B | HuggingFace |
| Wan2.1-T2V-1.3B | Text encoder, tokenizer, VAE and DiT architecture config | HuggingFace |
| DA3 (Depth-Anything-3) | Depth estimation (Step 1) | HuggingFace |
bash pipeline/download.shExpected directory structure after downloading:
checkpoints/
├── InSpatio-World-1.3B/
│ └── InSpatio-World-1.5-1.3B.safetensors
├── Wan2.1-T2V-1.3B/
└── depth/
Supported inputs are a single image, four images, or a video. Results are saved to
output/<output_id>/{source,render,mask,pred}.mp4.
For image and multi-image data, see examples/README.md.
bash run_example.shThis runs the six scenes listed in examples/manifest.json.
Image and multi-image scenes reuse uint16 depth PNGs. For video, DA3 estimates
both depth and per-frame source camera intrinsics/extrinsics when they are
missing. The target trajectory is still supplied by the user.
Use --depth_mode existing to require existing depth, or --depth_mode estimate
to regenerate it; the default is auto.
For a video, supply only the RGB video, prompt text, and one target OpenCV world-to-camera 4×4 matrix per frame (16 row-major numbers per line):
bash run_inference.sh --video /path/to/input.mp4 \
--prompt "Your prompt" --target_traj /path/to/target_tcw.txtThe runner reads the video's fps and frame count, resizes it to 832×480 if needed, and estimates depth and per-frame source cameras with DA3. The target trajectory must use the same first-frame-normalized coordinate system and displacement scale as the estimated source cameras.
This project is licensed under the Apache-2.0 License. Note that this license only applies to code in our library; dependencies such as Depth-Anything-3 are separately licensed.
If you use InSpatio-World in your research, please use the following BibTeX entry.
@misc{inspatio-world,
title={INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling},
author={InSpatio Team},
journal={arXiv preprint arXiv: 2604.07209},
year={2026}
}InSpatio-World utilizes a backbone based on Wan2.1, with its training code referencing Self-Forcing. We thank the Self-Forcing and Wan teams for their work and open-source contributions. We also acknowledge Depth-Anything-3 and ReCamMaster for their contributions.

