Skip to content
Oct 6Tue
  1. arXiv Game AI42

    SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games

    The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.

    Why it matters: The paper uses speedrun tasks across 9 games to probe agents' strategy formation, showing near-human world-record performance on simple platformers while still trailing on long, complex games.

  2. arXiv Game AI51

    PlaySuite: a large-scale benchmark for interactive visual intelligence covering more than 5K open-source games

    The authors propose PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence built on more than 5K open-source games from PyWeek and itch.io, spanning engines including Pygame, HTML5, Godot and Unity.

    Why it matters: The benchmark covers open-source games across multiple engines and provides a unified evaluation protocol, useful for comparing how different models sustain action in dynamic visual environments.