🤖 SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
⭐ If you find SpatialSpeak interesting, please consider starring this repository. Thank you!
Yang Cao1, Jiaxin Zhang3, Dave Zhenyu Chen2, Yingji Zhong1, Ruiyuan Gao2, Lanqing Hong2, Dan Xu*1
1 The Hong Kong University of Science and Technology
2 Huawei Noah’s Ark Lab
3 Harbin Institute of Technology
-
Our paper is now available on arXiv.
-
The code has not yet been released. Please stay tuned for updates.
Learning local geometry and global context makes spatial CoT more effective.
On ReVSI, reconstruction pretraining increases the gain from spatial CoT learning from 2.6 to 6.9 points. SpatialSpeak-4B achieves 62.8, exceeding the strongest compared baseline by 8.7 points.
A two-stage framework that connects reconstruction and reasoning through a shared text-based question-answering interface:
Visit our project page for the demo video, interactive point clouds, and additional examples.
If you find SpatialSpeak useful for your research, please consider citing:
@article{cao2026spatialspeak,
title={SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning},
author={Cao, Yang and Zhang, Jiaxin and Chen, Dave Zhenyu and Zhong, Yingji and Gao, Ruiyuan and Hong, Lanqing and Xu, Dan},
journal={arXiv preprint arXiv:2609.33616},
year={2026}
}For questions, please contact Yang Cao.
We sincerely thank the authors of the following projects for sharing their research and resources with the community:
GeoThinker, SpatialStack, VG-LLM, VLM-3R, Qwen3-VL, ReVSI, VSI-Bench, SPAR, Cambrian-S, etc