Can AI agents learn their way to the top? AAArena evaluates heuristic learning in a long-running game competition
The paper proposes the AAArena benchmark, using 12 real adversarial games and 1,920 archived human programs to evaluate agents' learning ability in long-running competitions. Among the evaluated model and tool configurations, Opus5.5 with Claude Code took 6 gold medals, and no configuration could take the remaining 6 human ladders.
Why it matters: AAArena organizes 12 adversarial games and 1,920 human programs into a competition-style evaluation, which can be used to observe the limits of how agents improve their strategies from limited samples.