SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games
The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.
Why it matters: The paper uses speedrun tasks across 9 games to probe agents' strategy formation, showing near-human world-record performance on simple platformers while still trailing on long, complex games.