Skip to content
TodayOct 9Fri2 items
  1. arXiv Game AI56

    MultiWorldBench: can independently controlled views describe one shared world

    The authors propose MultiWorldBench, a diagnostic Minecraft benchmark for testing whether independently controlled views in a multi-player world model stay consistent with one single persistent shared world; it contains 495 case configurations, seven task suites and ten capabilities.

    Why it matters: With 495 configurations, the paper compares three generative world models against the reference Engine GT on ten capabilities, and the score gaps show where the current weaknesses lie.

  2. arXiv Game AI49

    Can AI agents learn their way to the top? AAArena evaluates heuristic learning in a long-running game competition

    The paper proposes the AAArena benchmark, using 12 real adversarial games and 1,920 archived human programs to evaluate agents' learning ability in long-running competitions. Among the evaluated model and tool configurations, Opus5.5 with Claude Code took 6 gold medals, and no configuration could take the remaining 6 human ladders.

    Why it matters: AAArena organizes 12 adversarial games and 1,920 human programs into a competition-style evaluation, which can be used to observe the limits of how agents improve their strategies from limited samples.

Oct 7Wed
  1. arXiv Game AI47

    Paper proposes Recursive Game Creator, using four recursively iterating components to improve generated game experience

    The paper proposes Recursive Game Creator, an experience-oriented agentic game development framework that uses four components — Designer, Builder, Player and Reviewer — to iterate recursively and push a rough game prototype toward a more replayable work.

    Why it matters: The paper presents the four-component recursive workflow and results on two benchmarks, with 53.2% success rate on strict tasks in GameASG-Bench, a 34.1% improvement over the same-model baseline.

Oct 5Mon