GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, et al.
arXiv:2609.25001v1Today’s pick is GameHorizon Suite, a new benchmark and dataset for measuring how well AI systems can play video games across multiple time horizons. The problem it tackles is that existing game benchmarks are often narrow, noisy, or hard to reproduce: they may test only one kind of skill, lack language instructions, or depend on live rollouts that vary from run to run. GameHorizon builds a more complete testbed. It includes an automated annotation pipeline for instructions at different horizons, a large dataset of 5,000 hours of gameplay from 21 AAA games, and a benchmark with both offline and stepwise online evaluation. That means models are tested not just on what they can recognize, but on whether they can plan, act, and recover over long sequences of decisions. The authors evaluate 47 models and find clear differences in what kinds of gameplay abilities each one has. This matters because games are a compact way to study perception, planning, language, and control together.
Also spotted that day
Previous daily papers