← Back
TencentARC

TencentARC/GameHorizon

The official repo of GameHorizon, a unified data and evaluation suite that measures AAA gameplay capabilities at different temporal horizons for diverse model families.

View on GitHub ↗
Stars
442
Forks
1
Watchers
442
Open issues
1
Contributors
1
Language
—
License
Apache License 2.0
Default branch
main
Created Sep 20, 2026Updated Sep 22, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

GameHorizon Suite — Multi-Horizon Data and Evaluation in Gameplay


Yiran Wang1,*,†, Xingyilang Yin1,2,6,*, Junfu Pu1,*, Guangzhi Wang1,*, Kaifeng Li2, Mingyu Ouyang1,3, Huiqiang Sun1,4,
Lingen Li1,5, Cheng Cheng1, Wangbo Yu1, Honghao Chen1, Xiaodong Cun1,2,✉, Chi-Man Pun6, Zhiguo Cao4, Ying Shan1

1ARC Lab, Tencent   2GVC Lab, Great Bay University   3NUS   4HUST   5MMLab, CUHK   6University of Macau
† Project Lead    * Equal Contribution    ✉ Corresponding Author


📢 News

  • [2026.09.22] 📄 The paper and project page of GameHorizon Suite are released!
  • We will progressively release the code, benchmark, and data starting around 2026.10.25. Stay tuned! ⭐

🚀 Open-Source Plan

We are actively preparing the release of our code, benchmark, and data. Star the repo to follow our progress.

  • 🌐 Project page
  • 📄 Technical report / paper
  • 🏆 Public leaderboard
  • 🏗️ GameHorizon-Annotator — annotation toolkit & pipeline
  • 📊 GameHorizon-Bench (offline) — MCQ suite & evaluation code
  • 📊 GameHorizon-Bench (online) — online task environments
  • 🎯 GameHorizon-Data — 5,000-hour AAA gameplay corpus

📖 Overview

GameHorizon Suite overview

We introduce GameHorizon, a large-scale data and evaluation suite spanning multiple horizons and AAA games. It serves as a unified yardstick across a broad range of model types. It consists of three key components. GameHorizon-Annotator automatically produces a three-level pyramid of short-horizon operations, medium-horizon goals, and long-horizon strategies. GameHorizon-Data contains 5,000 hours of gameplay across 21 game titles, with temporally aligned videos, actions, and multi-horizon instructions. GameHorizon-Bench provides reproducible offline and stepwise online testing. The offline track contains thousands of standardized MCQs across three primary tasks and diagnostic variants. The online track tests order-dependent causal and order-flexible thematic tasks via the verifiable subtasks for failure localization.

Component What it is Highlights
🏗️ Annotator A scalable and automated annotation pipeline L1 → L2 → L3 instruction pyramid
🎯 Data A large-scale AAA gameplay corpus 5,000 hours · AAA-focused · 21 titles · temporally-aligned videos, actions & multi-horizon instructions
📊 Bench Reproducible offline + stepwise online track 5,000 offline MCQs (3 primary + 10 variant tasks) · 20 online tasks / 62 subtasks · 47 models benchmarked

🏆 Leaderboard

Offline Results — Primary Tasks (T1 / T2 / T3)

Accuracies are reported as percentages. Overall is the mean across the three tasks.

Bold = best, underline = second-best. Models are grouped into four tiers by Overall score.

Average (all 44 models) — T1 57.3 · T2 65.1 · T3 71.6 · Overall 64.7

🥇 Tier 1 · Rank 1–11

#ModelT1T2T3Overall
🥇GPT-6-Astra69.479.691.580.2
🥈Gemini 3.8 Flash65.981.284.877.3
🥉Gemini 3.7 Flash66.280.183.976.7
4Gemini 3.6 Flash62.379.184.675.3
5GPT-5.6 Sol65.377.581.574.8
6Kimi-K364.479.379.774.5
7Gemini 3.5 Flash64.977.480.874.4
8GPT-5.564.678.976.273.2
9Gemini 3.1 Pro64.575.878.272.8
10Doubao-Seed-2.1-Turbo63.179.475.572.7
11Doubao-Seed-2.1-Pro62.976.776.672.1

Tier 2 · Rank 12–22

#ModelT1T2T3Overall
12Doubao-Seed-2.0-Pro59.274.980.071.4
13Claude Fable 562.972.778.171.2
14GPT-5.6 Terra62.175.674.470.7
15Qwen3.7-Plus58.875.971.568.7
16Qwen3.8-27B59.977.468.168.5
17GPT-5.6 Luna61.670.771.167.8
17Kimi-K2.661.273.269.067.8
19Doubao-Seed-2.0-Lite55.269.076.366.8
20Gemma 4 31B-IT53.868.676.666.3
21MiniMax-M357.868.570.665.6
22Qwen3-VL-235B-A22B-Thinking58.271.665.965.2

Tier 3 · Rank 23–33

#ModelT1T2T3Overall
23Claude Opus 4.859.162.173.164.8
24Claude Sonnet 554.457.782.064.7
25Step-3.7-Flash57.571.363.464.1
26Qwen3-VL-235B-A22B-Instruct55.162.173.163.4
27GLM-5V-Turbo53.465.470.863.2
28GPT-5.258.958.970.662.8
29Qwen3.5-397B-A17B53.360.173.562.3
30BAGEL-7B-MoT51.757.375.161.4
31Qwen2.5-VL-32B-Instruct54.858.069.760.8
32Step3-VL-10B55.464.761.760.6
33SenseNova-U1-8B-MoT48.259.267.558.3

Tier 4 · Rank 34–44

#ModelT1T2T3Overall
34Qwen3.6-35B-A3B53.352.966.357.5
35Qwen3-Omni-30B-A3B-Instruct52.555.663.757.3
36GELab-Zero-4B-Preview51.846.872.757.1
37Qwen2.5-VL-7B-Instruct52.054.363.756.7
38InternVL3.5-8B51.850.265.555.8
39GLM-4.1V-9B-Thinking54.954.855.355.0
40GPT-4o55.746.361.354.4
41UI-TARS-1.5-7B49.147.162.953.0
42Ovis-U1-3B44.739.659.247.8
43InternVL3.5-2B45.140.354.746.0
44InternVL-U-4B46.137.949.844.6

Online Results — Long Horizon Tasks and Short-Horizon Subtasks

Each entry reports the success rate (%) with the number of passed tasks in parentheses.

The offline and online rankings show a clear positive association.

Online Model Offline Short-Horizon Subtasks Long-Horizon Tasks Causal Thematic
🥇 GPT-6-Astra 1 66.1 (41/62) 45.0 (9/20) 40.0 (4/10) 50.0 (5/10)
🥈 Gemini 3.6 Flash 4 56.5 (35/62) 30.0 (6/20) 30.0 (3/10) 30.0 (3/10)
🥉 Kimi-K3 6 46.8 (29/62) 10.0 (2/20) 20.0 (2/10) 0.0 (0/10)
4 GPT-5.6 Terra 14 37.1 (23/62) 10.0 (2/20) 20.0 (2/10) 0.0 (0/10)
5 GPT-5.6 Luna 17 33.9 (21/62) 10.0 (2/20) 10.0 (1/10) 10.0 (1/10)
6 MiniMax-M3 21 29.0 (18/62) 10.0 (2/20) 10.0 (1/10) 10.0 (1/10)
7 GLM-5V-Turbo 27 27.4 (17/62) 5.0 (1/20) 10.0 (1/10) 0.0 (0/10)
8 Qwen3.5-397B-A17B 29 19.4 (12/62) 5.0 (1/20) 10.0 (1/10) 0.0 (0/10)
9 Step3-VL-10B 32 14.5 (9/62) 5.0 (1/20) 10.0 (1/10) 0.0 (0/10)
10 Qwen3.6-35B-A3B 34 11.3 (7/62) 5.0 (1/20) 10.0 (1/10) 0.0 (0/10)
11 InternVL3.5-8B 38 3.2 (2/62) 0.0 (0/20) 0.0 (0/10) 0.0 (0/10)
12 UI-TARS-1.5-7B 41 1.6 (1/62) 0.0 (0/20) 0.0 (0/10) 0.0 (0/10)

📜 Citation

If you find GameHorizon Suite useful, please consider citing:

@article{GameHorizonSuite2026,
      title={GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay}, 
      author={Yiran Wang and Xingyilang Yin and Junfu Pu and Guangzhi Wang and Kaifeng Li and Mingyu Ouyang and Huiqiang Sun and Lingen Li and Cheng Cheng and Wangbo Yu and Honghao Chen and Xiaodong Cun and Chi-Man Pun and Zhiguo Cao and Ying Shan},
      year={2026},
      journal={arXiv preprint arXiv:2609.25001},
      url={https://arxiv.org/abs/2609.25001}, 
}

© 2026 ARC Lab, Tencent · GameHorizon Suite