Zhen Wang1 · Changpeng Wang1 · Zhe Liu2 · Zhangyang Qi2
Yuxiang Lu2 · Zimo Zeng1 · Donglian Qi1 · Xi Chen2
1Zhejiang University 2The University of Hong Kong
PanoVLN_main_1080p.mp4
Getting Started · Training Data · Dataset Construction · Models · Training · Evaluation · Robot Deployment · Configuration · Citation
PanoVLN is a vision-and-language navigation policy that follows instructions from 360° RGB observations. It combines a Qwen3.5-4B backbone with PanoVGGT geometry features and predicts 18-action sequences. Confidence-guided execution selects how far to move before observing and planning again.
This repository provides training and data-generation code, Habitat evaluation, three released checkpoints, and a GPU-server/Go2-client deployment pipeline. For physical navigation, PanoVLN Real World adds trajectory recovery and collision avoidance, helping the robot handle route deviations and obstacles during execution.
| Task | Entry point |
|---|---|
| Evaluate a released checkpoint | Environment → data layout → evaluation |
| Train or fine-tune | Training data → training |
| Deploy on a Unitree Go2 | Real-world deployment |
| Generate new trajectories and instructions | Dataset-generation guide |
| Read the method and full results | Paper on arXiv · project page |
- 2026-09-25: Source code released, including training, data generation, simulation evaluation, and robot deployment.
- Available: Three checkpoints, training annotations, and the project page with a real-world navigation video and main results.
- 2026-09-28: The paper is available on arXiv.
Training and model inference use Linux with an NVIDIA GPU. The model environment uses Python 3.12, PyTorch 2.10.0, and Transformers 5.5.0. Simulation additionally uses Habitat-Sim / Habitat-Lab 0.3.3; the Go2 client uses its separate ROS2 environment.
git clone https://github.com/wangzhen-w/PanoVLN.git
cd PanoVLN
conda create -n panovln python=3.12 -y
conda activate panovln
python -m pip install torch==2.10.0 torchvision==0.25.0
python -m pip install \
transformers==5.5.0 accelerate==1.13.0 peft==0.18.1 \
deepspeed==0.18.0 numpy pillow pyyaml tqdm scipy einops timm \
omegaconf huggingface-hub safetensors wandbChoose PyTorch CUDA wheels for your driver. Install FlashAttention 2 for the supplied launchers, or select sdpa in their attention settings. Follow the Habitat installation steps before rendering or evaluation. PanoVLN runs directly from the repository; its launchers set PYTHONPATH.
Download the simulation model:
hf download wangzhen-w/PanoVLN --local-dir checkpoints/PanoVLNKeep the checkpoint directory, including tokenizer, processor, and configuration files. See the reproduction guide for installation checks, exact configuration, outputs, and troubleshooting.
Download the PanoVLN training dataset from Hugging Face and follow the dataset repository instructions to prepare the released files.
hf download wangzhen-w/PanoVLN --repo-type dataset --local-dir dataThe download contains navigation annotations and training JSONL files. Obtain scene assets separately and render the required panoramic images.
For R2R-CE and RxR-CE, follow the VLN-CE dataset instructions. Obtain the corresponding Matterport3D scene assets separately. Creating additional PanoVLN trajectories requires HM3D scenes.
All dataset paths are relative to the repository root, with data/ as the default data root. Place your dataset there, or create a symlink at that location to an existing dataset directory. Arrange the downloaded files as follows; keep the split subdirectories in R2R, RxR, and HM3D:
data/
├── general_vln_dataset/
│ ├── r2r/
│ │ ├── train/
│ │ │ └── train.json.gz
│ │ ├── val_seen/
│ │ │ └── val_seen.json.gz
│ │ └── val_unseen/
│ │ └── val_unseen.json.gz
│ ├── rxr/
│ │ ├── train/
│ │ │ └── train_guide.json.gz
│ │ ├── val_seen/
│ │ │ └── val_seen_guide.json.gz
│ │ └── val_unseen/
│ │ └── val_unseen_guide.json.gz
│ └── panovln/
│ └── train.json.gz
├── scene/
│ ├── hm3d/
│ │ ├── train/
│ │ │ ├── 00000-kfPV7w3FaU5/
│ │ │ ├── 00001-UVdNNRcVyV1/
│ │ │ └── ...
│ │ └── val/
│ │ └── ...
│ └── mp3d/
│ ├── 17DRP5sb8fy/
│ ├── 1LXtFkjw3qL/
│ └── ...
├── images/
│ ├── r2r/
│ │ └── <episode_id>/
│ │ └── frame_0.jpg
│ ├── rxr/
│ │ └── <episode_id>/
│ │ └── frame_0.jpg
│ ├── panovln/
│ │ └── <episode_id>/
│ │ └── frame_0.jpg
│ └── dagger/
│ └── <trajectory_id>/
│ └── frame_0.jpg
├── sub_dataset/
│ ├── r2r.jsonl
│ ├── rxr.jsonl
│ ├── panovln.jsonl
│ └── dagger.jsonl
├── train_r2r_rxr.jsonl # R2R + RxR mixture
├── r2r_rxr_dagger.jsonl # R2R + RxR + DAgger mixture
├── r2r_rxr_dagger_panovln.jsonl # Full mixture
└── train.jsonl # Mixture selected in the training config
Only the datasets used in your run are required. Keep the original scene assets and their navigation meshes. The paths in config/ follow this layout.
general_vln_dataset/ holds navigation episodes; scene/ holds simulator scene assets. sub_dataset/ holds action-aligned annotations, and images/ holds their panoramic observations. The JSONL files at the data root are prepared training mixtures. Only evaluation episodes and scenes are needed to evaluate a checkpoint.
Data scripts process R2R and RxR by default. Change DATASET_NAMES at the top of each script to select other datasets (SOURCE_DATASET_NAMES for DAgger collection).
Run preprocessing and frame extraction in order:
bash scripts/preprocess.sh
bash scripts/extract_frame.shAnnotations are saved to data/sub_dataset/, and images to data/images/. Skip the corresponding step if these files are already prepared.
Build the training JSONL with scripts/prepare_dataset.sh:
bash scripts/prepare_dataset.shThe default output is data/train.jsonl, containing 18-action training samples. If you use a released mixture, select it directly in data.train_jsonl in src/train/config/config.yaml:
| Training stage | JSONL file | Sources |
|---|---|---|
| Initial policy | data/train_r2r_rxr.jsonl |
R2R-CE + RxR-CE |
| PanoVLN Base (†), including DAgger refinement | data/r2r_rxr_dagger.jsonl |
R2R-CE + RxR-CE + corrective trajectories |
| Full PanoVLN | data/r2r_rxr_dagger_panovln.jsonl |
The above + the constructed PanoVLN dataset |
| Custom mixture | data/train.jsonl |
Generated from DATASET_NAMES in scripts/prepare_dataset.sh |
Set data.train_image_root to data/: sample image paths already start with images/. Render all image sources selected by the mixture. See the training-data preparation guide for regeneration and DAgger.
To create new trajectories and language instructions from HM3D scenes, use dataset_create/. Its guide covers scene inspection, trajectory collection and replay validation, instruction generation, and export of panoramic training images. This is separate from preparing or selecting the released training mixtures above.
| Model | Role | Repository / weights | Suggested local path |
|---|---|---|---|
| PanoVLN | Full-data checkpoint for simulation benchmarks | Hugging Face | checkpoints/PanoVLN/ |
| PanoVLN Base (†) | R2R/RxR-only navigation-data checkpoint | Hugging Face | checkpoints/PanoVLN_base/ |
| PanoVLN Real World | Robot deployment with enhanced trajectory recovery and collision avoidance | Hugging Face | checkpoints/PanoVLN_realworld/ |
| Qwen3.5-4B | VLM initialization for training | Hugging Face | checkpoints/Qwen3.5-4B/ |
| PanoVGGT | Frozen panoramic geometry encoder | Upstream checkpoint | checkpoints/PanoVGGT/model.pt |
Use PanoVLN to reproduce the full-data simulation results and PanoVLN_base for the PanoVLN† results.
For physical robot deployment, we provide a dedicated PanoVLN_realworld checkpoint with two enhancements:
- Trajectory recovery: Improved ability to recover from deviations and resume following the navigation instruction.
- Collision avoidance: Improved ability to avoid obstacles during navigation.
These enhancements address recovery and collision avoidance during physical execution, which motivates a separate deployment checkpoint. Use PanoVLN_realworld with the real-world deployment guide; use the simulation checkpoints above to reproduce the reported benchmark results.
Use a fully saved PanoVLN checkpoint, including its tokenizer and processor files, for evaluation or deployment.
hf download wangzhen-w/PanoVLN --local-dir checkpoints/PanoVLN
# Alternatives: PanoVLN_base or PanoVLN_realworld, with the matching local directory.Main benchmark results are shown on the project page.
hf download Qwen/Qwen3.5-4B --local-dir checkpoints/Qwen3.5-4B
hf download YijingGuo/PanoVGGT --local-dir checkpoints/PanoVGGTConfigure the following before launching:
| File | Settings |
|---|---|
src/train/config/config.yaml |
model.name_or_path, model.panovggt_checkpoint_path, data.train_jsonl, data.train_image_root, batch size, attention, and learning rates |
src/train/train.sh |
GPU_DEVICES, CONFIG_PATH, and a fresh OUTPUT_DIR |
Defaults use checkpoints/Qwen3.5-4B/, checkpoints/PanoVGGT/model.pt, data/train.jsonl, and data/. The launcher defaults to eight GPUs; set the GPU list and batch size for your machine. The PanoVGGT encoder remains frozen while the selected language, visual, and fusion modules are trained.
bash src/train/train.shOUTPUT_DIR receives train.log, periodic checkpoint-* directories, and the final model/tokenizer/processor. Validation is disabled by default. The launcher takes its settings from the files above and rejects command-line overrides; see the training guide for batch-size accounting, validation, and resuming.
Set the trained MODEL_PATH, source datasets, GPUs, and output paths in scripts/generate_dagger_data.sh, then collect corrective trajectories:
bash scripts/generate_dagger_data.shInclude dagger in DATASET_NAMES in scripts/prepare_dataset.sh, point model.name_or_path to the starting navigation checkpoint, and use a new training output directory:
bash scripts/prepare_dataset.sh --overwrite
bash src/train/train.shThe reproduction guide explains the generated annotations, images, and mixture selection.
Edit scripts/eval_r2r.sh or scripts/eval_rxr.sh. For a first R2R run, set these variables inside the launcher:
MODEL_PATH="./checkpoints/PanoVLN"
GPU_IDS="0"
PROCS_PER_GPU=1
TOTAL_MAX_EPISODES=5
SAVE_PATH="./outputs/eval/r2r_smoke"Then launch the selected benchmark:
bash scripts/eval_r2r.sh
# For RxR, configure the other launcher and use a separate SAVE_PATH:
# bash scripts/eval_rxr.shSet TOTAL_MAX_EPISODES=0 for the complete split. Both launchers default to val_unseen, uncertainty-based execution, and a 4–8-action range. Preserve their benchmark-specific stop settings when reproducing reported results. Each worker loads a model; adjust PROCS_PER_GPU to available memory.
| Output | Contents |
|---|---|
result.jsonl |
Per-episode metrics and predicted/executed action histories |
result_summary.json |
Episode count, success, SPL, oracle success, distance to goal, path length, and nDTW |
top_down/ |
Per-episode videos when SAVE_TOPDOWN=true |
Existing episodes are skipped on resume. Use a new SAVE_PATH for each checkpoint or policy configuration. See evaluation settings for single-worker commands and output interpretation.
Use PanoVLN_realworld with a GPU inference server and a panoramic-camera Unitree Go2 client. This checkpoint adds trajectory recovery after route deviations and collision avoidance during physical execution.
Follow realworld/ for hardware and environment setup, checkpoint configuration, server startup, ROS2 client setup, navigation and trial recording. The PanoVLN reference documents configuration fields and the prediction API.
| Change | Edit |
|---|---|
| Checkpoint, data mixture, trainable modules, precision, learning rates | src/train/config/config.yaml |
| Training GPUs and output directory | src/train/train.sh |
| Evaluation workers, memory window, and execution policy | scripts/eval_r2r.sh, scripts/eval_rxr.sh |
| Dataset split, scene paths, panoramic sensor | config/ |
| Preprocessing, rendering, and custom data mixtures | DATASET_NAMES in scripts/ |
| Robot server address, camera, instruction, motion, recording | realworld/panovln/go2_client.yaml |
Action targets contain 18 words from forward (0.25 m), left (15°), right (15°), and stop; terminal targets are padded with stop. actions_per_replan controls the executed prefix independently. The training configuration guide documents the architecture constraints, ERP cropping, and fusion options.
If you use PanoVLN, please cite our paper.
@misc{wang2026panovln,
title = {PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation},
author = {Zhen Wang and Changpeng Wang and Zhe Liu and Zhangyang Qi and
Yuxiang Lu and Zimo Zeng and Donglian Qi and Xi Chen},
year = {2026},
eprint = {2609.34759},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2609.34759},
url = {https://arxiv.org/abs/2609.34759}
}Machine-readable citation metadata is available in CITATION.cff.
Video music credits are listed in assets/README.md.
The repository-wide code license is pending an author decision. Bundled third-party components retain their PanoVGGT, NaVid, and NaVILA license notices.
The released Hugging Face model and dataset cards specify Matterport Academic Use terms. Consult each resource's card and the Matterport academic-use agreement before use. Obtain source scenes under their respective access and license terms.
PanoVLN builds on Qwen, PanoVGGT, Habitat, and VLN-CE. We thank their authors and the dataset contributors for making these resources available.