David Nordström1, Thibaut Loiseau2, Vincent Lepetit2,
Michael Felsberg3, Guillaume Bourmaud4, Fredrik Kahl1
1 Chalmers University of Technology, 2 Ecole Nationale des Ponts et Chaussées, IP Paris,
3 Linköping University, 4 University of Bordeaux, CNRS
Left: Emerging matching capabilities from self-supervision. Simply using Poincar3's attention map reliable tracks can be created. Produced by demo.py. Right: Picture of Henri Poincaré, who argued that a being without motion cannot understand 3D space. Motivating our use of multi-view data.
We release a 3D foundation model, Poincar3, that learns multi-view geometry from only training on image sequences. Poincar3 uses a multi-view transformer and self-distillation, achieving strong zero-shot features and attention maps. For example, you can finetune our model for just 10K steps on one GPU and get 65+ AUC@30 on RE10K, whereas training from scratch gives around 5 AUC@30.
- [September 30, 2026] Poincar3 public code release.
uv add poincar3 # or: pip install poincar3The model itself needs only torch. Extras pull in what the scripts need:
uv add "poincar3[demo]" # demo.py
uv add "poincar3[train]" # training, mvcorr and SeeSE3 evals
uv add "poincar3[eval]" # + matchbench and feed-forward reconstructionTo work from a clone instead (tested on Linux with Python 3.12):
uv sync --extra evalThe checkpoint auto-downloads on first use:
import torch
from poincar3 import Poincar3
model = Poincar3().eval().cuda()
# A multi-view batch: [batch, frames, 3, H, W], RGB in [0, 1], H and W multiples of 16.
images = torch.rand(1, 4, 3, 448, 448).cuda()
with torch.no_grad():
patch_logits, patch_features, global_logits, camera_tokens = model(images)
# patch_features: [1, 4, 784, 1024] dense per-frame tokens, cross-view attended
# camera_tokens: [1, 4, 1024] one per-frame scene/pose tokenWe illustrate how to create attention tracks by simply running:
uv run python demo.pyWe also provide code for plotting the raw feature correlations in cross_correlations.py.
The backbone auto-downloads on first use. You can find it directly here.
All evaluations, except feed-forward reconstruction, can be accessed through the experiments/eval.py endpoint. You can get ScanNet and NAVI following these instructions. To get the possible configurations, simply run it with the flag --help. For example, you run multi-view correspondence estimation on scannet by:
uv run python experiments/eval.py --evaluation mvcorr --mvcorr.dataset scannetThis should give an accuracy at 50px of around 89.6 whereas using --mvcorr.correspondence-method attention should instead give around 94.9. We also give easy access to baselines by using e.g. --backbone dinov3_vitl.
Camera-pose and per-pixel-depth heads on top of the backbone, under two protocols (we train on 4xH200):
# full finetune
torchrun --nproc-per-node 4 experiments/ffrecon/train.py --name finetune-poincar3
# frozen backbone + a small adapter
torchrun --nproc-per-node 4 experiments/ffrecon/train_adapter.py --backbone poincar3
# relative pose evaluation
uv run python experiments/ffrecon/eval.py --checkpoint <path> --evaluation relpose --relpose.dataset megadepth
# point-cloud estimation
uv run python experiments/ffrecon/eval.py --checkpoint <path> --evaluation pointcloud --pointcloud.dataset eth3dWe also provide the pretrained checkpoints directly, you can use by experiments/ffrecon/eval.py --checkpoint <file>, and find them here:
| checkpoint | protocol | backbone | RE10K AUC@30 |
|---|---|---|---|
poincar3_ffrecon_finetune.pth |
full finetune | Poincar3 | 0.682 |
poincar3_ffrecon_finetune_dinov3_init.pth |
full finetune | DINOv3 init | 0.312 |
poincar3_ffrecon_finetune_random_init.pth |
full finetune | random init | 0.174 |
poincar3_ffrecon_adapter.pth |
frozen + adapter | Poincar3 | 0.622 |
dinov3_ffrecon_adapter.pth |
frozen + adapter | DINOv3 | 0.291 |
mum_v1_ffrecon_adapter.pth |
frozen + adapter | MuM v1 | 0.401 |
We pretrain Poincar3 on 8xH200 for 3 days. We run the command:
torchrun --nproc-per-node 8 experiments/train.py --name my-runWhile we trained on large collection of 3D datasets, we illustrate our training protocol on ScanNet++ and RealEstate10K. You can download them using their official download links.
MIT, except where a file notes otherwise. src/poincar3/layers/ and parts of heads/ derive
from DINOv3 and
VGGT and carry their original licenses;
benchmarks/mv_consistency/ is adapted from probe3d (MIT).
Built on DINOv3, DINOv2, VGGT, probe3d and RoMa.
What data/compute did you use?
We pretrained our 650M multi-view transformer on 8xH200 for 3 days. The following datasets were used:
| Datasets | Type / Source | Weight | # Scenes |
|---|---|---|---|
| SpatialVID | Outdoor / Video | 1 | 176,749 |
| DL3DV | Mixed / Video | 1 | 10,000 |
| RealEstate10K | Indoor / Video | 1 | 7,850 |
| MegaDepth | Outdoor / MVS | 1 | 169 |
| AerialMD | Aerial / MVS | 1 | 124 |
| BlendedMVS | Aerial / Mesh | 1 | 493 |
| Hypersim | Indoor / Graphics | 1 | 393 |
| TartanAir v2 | Outdoor / Graphics | 1 | 46 |
| Map-Free | Object-centric / MVS | 1 | 397 |
| ScanNet++ v2 | Indoor / Mesh | 1 | 856 |
| FlyingThings3D | Outdoor / Graphics | 0.5 | 2,239 |
| ARKitScenes | Indoor / RGB-D | 0.1 | 5,047 |
| UnrealStereo4k | Outdoor / Graphics | 0.01 | 8 |
| Virtual KITTI 2 | Outdoor / Graphics | 0.01 | 5 |
| Total | 204,376 |
Can you show me training curves?
You can find full evaluation and training curves in the Appendix of the paper.
How do you prevent collapse?
Similarly to DINO, wew use Sinkhorn-Knopp and Koleo regularization. These ensure that the output distribution cannot be uni-modal over a batch, ensuring diversity of clusters and preventing collapse (see CAPI paper for a good explanation).
Please feel free to email me directly at davnords@chalmers.se for any questions.
@misc{nordstrom2026emergentmultiview,
title={Emergent Multi-View Geometry Through Self-Distillation},
author={David Nordström and Thibaut Loiseau and Vincent Lepetit and Michael Felsberg and Guillaume Bourmaud and Fredrik Kahl},
year={2026},
eprint={2609.39227},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.39227},
}