Skip to content
Hot storyWatching

Researchers propose SPEEDRUNBENCH, a speedrunning benchmark spanning 9 games

1 report1 sourceUpdated 3 d ago

What happened

AI overview

Researcher Yoshinari Fujinuma proposes the speedrunning benchmark SPEEDRUNBENCH, which evaluates frontier LLM agents' strategy formation across 9 different games: agents must repeatedly improve their strategy over long-horizon actions, review their own performance and make use of knowledge they have already gained, to be faster than both their past selves and others. Experiments show that frontier agents already come close to human world records on simple platformers, but still lag behind human performance on longer, more complex games due to real budget constraints. The authors argue that the benchmark can serve long-term as a testbed for studying agents' strategy formation.

AI-written from the coverage · updated 2 h ago

Coverage timeline

Follow the coverage to see the story from different sides.

Oct 6
  1. arXiv Game AIPick
    SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games

    The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.

Heat trend for this story

Not enough continuous observations to chart a trend yet.