Skip to content
  1. arXiv Game AI40

    Mine Odyssey proposes an agentic spatial intelligence benchmark with 180 tasks, rebuilding 30 real locations in Minecraft

    Yuxuan Cao and others propose Mine Odyssey, an agentic spatial intelligence benchmark that rebuilds 30 real locations across 20 countries and regions on five continents in Minecraft, with 180 natural language tasks, 20 of them outdoor scenes and 10 indoor scenes.

    Why it matters: The paper rebuilds 30 real locations in Minecraft and sets 180 tasks, making it possible to compare the spatial intelligence gaps across eight models.

  2. arXiv Game AI56

    MultiWorldBench: can independently controlled views describe one shared world

    The authors propose MultiWorldBench, a diagnostic Minecraft benchmark for testing whether independently controlled views in a multi-player world model stay consistent with one single persistent shared world; it contains 495 case configurations, seven task suites and ten capabilities.

    Why it matters: With 495 configurations, the paper compares three generative world models against the reference Engine GT on ten capabilities, and the score gaps show where the current weaknesses lie.

  3. arXiv Game AI49

    Can AI agents learn their way to the top? AAArena evaluates heuristic learning in a long-running game competition

    The paper proposes the AAArena benchmark, using 12 real adversarial games and 1,920 archived human programs to evaluate agents' learning ability in long-running competitions. Among the evaluated model and tool configurations, Opus5.5 with Claude Code took 6 gold medals, and no configuration could take the remaining 6 human ladders.

    Why it matters: AAArena organizes 12 adversarial games and 1,920 human programs into a competition-style evaluation, which can be used to observe the limits of how agents improve their strategies from limited samples.

  4. arXiv Game AI45

    AgentGarten: a code world framework for continually evolving agents

    The paper proposes the AgentGarten framework, which couples simulators and game engines to a shared neural renderer to build real-time interactive virtual environments. The simulation backend maintains persistent world state and executes program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a unified interface; the neural renderer is adapted from a pretrained video model to take geometry-conditioned input, is distilled with the proposed Adversarial Forcing, and has its inference optimized for real-time interaction.

    Why it matters: The paper wires simulators and game engines into a single neural renderer, and reports learning-efficiency results comparing an agent's 4 rounds of experience with millions of rounds of reinforcement learning.

  1. arXiv Game AI45

    Kuration SDK: Addressing the Virtual2Real Gap in World Models via Data Curation

    The paper proposes and open-sources along with it Kuration SDK, a general-purpose physical AI data curation toolkit that uses data curation to address the Virtual2Real gap in training action-conditioned world models. The authors train and evaluate diffusion world models on CounterStrike gameplay data, confirming that metrics such as FVD, LPIPS and JEDi do not correspond to qualitative playability, and argue that curating raw gameplay data before training begins and measuring multiple diagnostic properties is a more reliable signal.

    Why it matters: The paper trains diffusion world models on CounterStrike gameplay data, points out that visual similarity metrics such as FVD and LPIPS do not correspond to playability, and puts forward a data curation approach.

  2. arXiv Game AI47

    Paper proposes Recursive Game Creator, using four recursively iterating components to improve generated game experience

    The paper proposes Recursive Game Creator, an experience-oriented agentic game development framework that uses four components — Designer, Builder, Player and Reviewer — to iterate recursively and push a rough game prototype toward a more replayable work.

    Why it matters: The paper presents the four-component recursive workflow and results on two benchmarks, with 53.2% success rate on strict tasks in GameASG-Bench, a 34.1% improvement over the same-model baseline.

  3. arXiv Game AI36

    Study proposes ResNet-BiLSTM for multi-label perceptual bug detection from gameplay footage

    The study proposes a ResNet-BiLSTM model that performs multi-label perceptual bug detection on gameplay footage, achieving an F1 score of 85.78% on a benchmark dataset, and compares it with video classification models including Inflated 3D ConvNet and 3D ResNet.

    Why it matters: The paper reports an F1 of 85.78% for ResNet-BiLSTM and compares it with video classification models such as Inflated 3D ConvNet on the same task.

  1. arXiv Game AI42

    SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games

    The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.

    Why it matters: The paper uses speedrun tasks across 9 games to probe agents' strategy formation, showing near-human world-record performance on simple platformers while still trailing on long, complex games.

  2. arXiv Game AI41

    Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

    The researchers propose Attacca, which trains a vision goal-conditioned policy on complete search-to-interaction trajectories so long-horizon embodied agents can carry on to the next task from the position, orientation and world state left by the previous one. The method uses context-decoupled goal sampling, pairing each demonstration with a class-compatible masked goal image from another world, and provides supervision beyond action imitation through a goal mask prediction head, while introducing behavior phase conditioning that distinguishes the Search, Approach and Interact phases.

    Why it matters: The paper brings state-continuous long-horizon execution into the evaluation and reports success rate and completion comparisons against the strongest baseline.

  3. arXiv Game AI35

    Emoception proposes the SALFT framework, selectively fine-tuning a video ViT to recognize changes in player arousal

    The paper proposes SALFT (Selective Affective Layer Fine-Tuning), a framework that fine-tunes a Video Vision Transformer using a selection criterion based on layer-wise parameter L2 norm change to recognize changes in player arousal from gameplay footage.

    Why it matters: The paper uses layer-wise parameter L2 norm change as its selection criterion and shows that updating only about 8% of parameters approaches full fine-tuning, useful for comparing compute costs against your own player emotion recognition pipeline.

  4. arXiv Game AI51

    PlaySuite: a large-scale benchmark for interactive visual intelligence covering more than 5K open-source games

    The authors propose PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence built on more than 5K open-source games from PyWeek and itch.io, spanning engines including Pygame, HTML5, Godot and Unity.

    Why it matters: The benchmark covers open-source games across multiple engines and provides a unified evaluation protocol, useful for comparing how different models sustain action in dynamic visual environments.

  1. arXiv Game AI39

    Code2Games: enabling coding agents to generate playable game worlds from natural language

    Wei Wu proposes Code2Games, an agent framework that lets coding agents build structured game worlds on top of a Blender world generated from the same game intent. The framework coordinates scene analysis, gameplay planning, constrained game world generation and engine customization through shared scene and gameplay representations, and when adapting to Unreal Engine 5 uses compilation diagnostics, runtime feedback and playtest results for execution-guided refactoring.

    Why it matters: The paper coordinates the generation stages with shared scene and gameplay representations and compares generation quality on a benchmark containing ten fixed game prompts, making it easy to measure against existing approaches.

You've reached the end