Yiran Wang1,*,†, Xingyilang Yin1,2,6,*, Junfu Pu1,*, Guangzhi Wang1,*, Kaifeng Li2,
Mingyu Ouyang1,3, Huiqiang Sun1,4,
Lingen Li1,5, Cheng Cheng1, Wangbo Yu1,
Honghao Chen1, Xiaodong Cun1,2,✉, Chi-Man Pun6, Zhiguo Cao4, Ying Shan1
1ARC Lab, Tencent 2GVC Lab, Great Bay University 3NUS 4HUST 5MMLab, CUHK 6University of Macau
† Project Lead * Equal Contribution ✉ Corresponding Author
- [2026.09.22] 📄 The paper and project page of GameHorizon Suite are released!
- We will progressively release the code, benchmark, and data starting around 2026.10.25. Stay tuned! ⭐
We are actively preparing the release of our code, benchmark, and data. Star the repo to follow our progress.
- 🌐 Project page
- 📄 Technical report / paper
- 🏆 Public leaderboard
- 🏗️ GameHorizon-Annotator — annotation toolkit & pipeline
- 📊 GameHorizon-Bench (offline) — MCQ suite & evaluation code
- 📊 GameHorizon-Bench (online) — online task environments
- 🎯 GameHorizon-Data — 5,000-hour AAA gameplay corpus
We introduce GameHorizon, a large-scale data and evaluation suite spanning multiple horizons and AAA games. It serves as a unified yardstick across a broad range of model types. It consists of three key components. GameHorizon-Annotator automatically produces a three-level pyramid of short-horizon operations, medium-horizon goals, and long-horizon strategies. GameHorizon-Data contains 5,000 hours of gameplay across 21 game titles, with temporally aligned videos, actions, and multi-horizon instructions. GameHorizon-Bench provides reproducible offline and stepwise online testing. The offline track contains thousands of standardized MCQs across three primary tasks and diagnostic variants. The online track tests order-dependent causal and order-flexible thematic tasks via the verifiable subtasks for failure localization.
| Component | What it is | Highlights |
|---|---|---|
| 🏗️ Annotator | A scalable and automated annotation pipeline | L1 → L2 → L3 instruction pyramid |
| 🎯 Data | A large-scale AAA gameplay corpus | 5,000 hours · AAA-focused · 21 titles · temporally-aligned videos, actions & multi-horizon instructions |
| 📊 Bench | Reproducible offline + stepwise online track | 5,000 offline MCQs (3 primary + 10 variant tasks) · 20 online tasks / 62 subtasks · 47 models benchmarked |
Accuracies are reported as percentages. Overall is the mean across the three tasks.
Bold = best, underline = second-best. Models are grouped into four tiers by Overall score.
|
Tier 3 · Rank 23–33
|
Tier 4 · Rank 34–44
|
Each entry reports the success rate (%) with the number of passed tasks in parentheses.
The offline and online rankings show a clear positive association.
| Online | Model | Offline | Short-Horizon Subtasks | Long-Horizon Tasks | Causal | Thematic |
|---|---|---|---|---|---|---|
| 🥇 | GPT-6-Astra | 1 | 66.1 (41/62) | 45.0 (9/20) | 40.0 (4/10) | 50.0 (5/10) |
| 🥈 | Gemini 3.6 Flash | 4 | 56.5 (35/62) | 30.0 (6/20) | 30.0 (3/10) | 30.0 (3/10) |
| 🥉 | Kimi-K3 | 6 | 46.8 (29/62) | 10.0 (2/20) | 20.0 (2/10) | 0.0 (0/10) |
| 4 | GPT-5.6 Terra | 14 | 37.1 (23/62) | 10.0 (2/20) | 20.0 (2/10) | 0.0 (0/10) |
| 5 | GPT-5.6 Luna | 17 | 33.9 (21/62) | 10.0 (2/20) | 10.0 (1/10) | 10.0 (1/10) |
| 6 | MiniMax-M3 | 21 | 29.0 (18/62) | 10.0 (2/20) | 10.0 (1/10) | 10.0 (1/10) |
| 7 | GLM-5V-Turbo | 27 | 27.4 (17/62) | 5.0 (1/20) | 10.0 (1/10) | 0.0 (0/10) |
| 8 | Qwen3.5-397B-A17B | 29 | 19.4 (12/62) | 5.0 (1/20) | 10.0 (1/10) | 0.0 (0/10) |
| 9 | Step3-VL-10B | 32 | 14.5 (9/62) | 5.0 (1/20) | 10.0 (1/10) | 0.0 (0/10) |
| 10 | Qwen3.6-35B-A3B | 34 | 11.3 (7/62) | 5.0 (1/20) | 10.0 (1/10) | 0.0 (0/10) |
| 11 | InternVL3.5-8B | 38 | 3.2 (2/62) | 0.0 (0/20) | 0.0 (0/10) | 0.0 (0/10) |
| 12 | UI-TARS-1.5-7B | 41 | 1.6 (1/62) | 0.0 (0/20) | 0.0 (0/10) | 0.0 (0/10) |
If you find GameHorizon Suite useful, please consider citing:
@article{GameHorizonSuite2026,
title={GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay},
author={Yiran Wang and Xingyilang Yin and Junfu Pu and Guangzhi Wang and Kaifeng Li and Mingyu Ouyang and Huiqiang Sun and Lingen Li and Cheng Cheng and Wangbo Yu and Honghao Chen and Xiaodong Cun and Chi-Man Pun and Zhiguo Cao and Ying Shan},
year={2026},
journal={arXiv preprint arXiv:2609.25001},
url={https://arxiv.org/abs/2609.25001},
}