Skip to content
TodayOct 9Fri3 items
  1. arXiv Game AI40

    Mine Odyssey proposes an agentic spatial intelligence benchmark with 180 tasks, rebuilding 30 real locations in Minecraft

    Yuxuan Cao and others propose Mine Odyssey, an agentic spatial intelligence benchmark that rebuilds 30 real locations across 20 countries and regions on five continents in Minecraft, with 180 natural language tasks, 20 of them outdoor scenes and 10 indoor scenes.

    Why it matters: The paper rebuilds 30 real locations in Minecraft and sets 180 tasks, making it possible to compare the spatial intelligence gaps across eight models.

  2. arXiv Game AI49

    Can AI agents learn their way to the top? AAArena evaluates heuristic learning in a long-running game competition

    The paper proposes the AAArena benchmark, using 12 real adversarial games and 1,920 archived human programs to evaluate agents' learning ability in long-running competitions. Among the evaluated model and tool configurations, Opus5.5 with Claude Code took 6 gold medals, and no configuration could take the remaining 6 human ladders.

    Why it matters: AAArena organizes 12 adversarial games and 1,920 human programs into a competition-style evaluation, which can be used to observe the limits of how agents improve their strategies from limited samples.

  3. arXiv Game AI45

    AgentGarten: a code world framework for continually evolving agents

    The paper proposes the AgentGarten framework, which couples simulators and game engines to a shared neural renderer to build real-time interactive virtual environments. The simulation backend maintains persistent world state and executes program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a unified interface; the neural renderer is adapted from a pretrained video model to take geometry-conditioned input, is distilled with the proposed Adversarial Forcing, and has its inference optimized for real-time interaction.

    Why it matters: The paper wires simulators and game engines into a single neural renderer, and reports learning-efficiency results comparing an agent's 4 rounds of experience with millions of rounds of reinforcement learning.

Oct 7Wed
  1. arXiv Game AI45

    Kuration SDK: Addressing the Virtual2Real Gap in World Models via Data Curation

    The paper proposes and open-sources along with it Kuration SDK, a general-purpose physical AI data curation toolkit that uses data curation to address the Virtual2Real gap in training action-conditioned world models. The authors train and evaluate diffusion world models on CounterStrike gameplay data, confirming that metrics such as FVD, LPIPS and JEDi do not correspond to qualitative playability, and argue that curating raw gameplay data before training begins and measuring multiple diagnostic properties is a more reliable signal.

    Why it matters: The paper trains diffusion world models on CounterStrike gameplay data, points out that visual similarity metrics such as FVD and LPIPS do not correspond to playability, and puts forward a data curation approach.

  2. arXiv Game AI47

    Paper proposes Recursive Game Creator, using four recursively iterating components to improve generated game experience

    The paper proposes Recursive Game Creator, an experience-oriented agentic game development framework that uses four components — Designer, Builder, Player and Reviewer — to iterate recursively and push a rough game prototype toward a more replayable work.

    Why it matters: The paper presents the four-component recursive workflow and results on two benchmarks, with 53.2% success rate on strict tasks in GameASG-Bench, a 34.1% improvement over the same-model baseline.

Oct 6Tue
  1. arXiv Game AI42

    SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games

    The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.

    Why it matters: The paper uses speedrun tasks across 9 games to probe agents' strategy formation, showing near-human world-record performance on simple platformers while still trailing on long, complex games.

  2. arXiv Game AI41

    Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

    The researchers propose Attacca, which trains a vision goal-conditioned policy on complete search-to-interaction trajectories so long-horizon embodied agents can carry on to the next task from the position, orientation and world state left by the previous one. The method uses context-decoupled goal sampling, pairing each demonstration with a class-compatible masked goal image from another world, and provides supervision beyond action imitation through a goal mask prediction head, while introducing behavior phase conditioning that distinguishes the Search, Approach and Interact phases.

    Why it matters: The paper brings state-continuous long-horizon execution into the evaluation and reports success rate and completion comparisons against the strongest baseline.

  3. arXiv Game AI51

    PlaySuite: a large-scale benchmark for interactive visual intelligence covering more than 5K open-source games

    The authors propose PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence built on more than 5K open-source games from PyWeek and itch.io, spanning engines including Pygame, HTML5, Godot and Unity.

    Why it matters: The benchmark covers open-source games across multiple engines and provides a unified evaluation protocol, useful for comparing how different models sustain action in dynamic visual environments.

Oct 5Mon
Oct 4Sun
  1. arXiv Game AI39

    Code2Games: enabling coding agents to generate playable game worlds from natural language

    Wei Wu proposes Code2Games, an agent framework that lets coding agents build structured game worlds on top of a Blender world generated from the same game intent. The framework coordinates scene analysis, gameplay planning, constrained game world generation and engine customization through shared scene and gameplay representations, and when adapting to Unreal Engine 5 uses compilation diagnostics, runtime feedback and playtest results for execution-guided refactoring.

    Why it matters: The paper coordinates the generation stages with shared scene and gameplay representations and compares generation quality on a benchmark containing ten fixed game prompts, making it easy to measure against existing approaches.