Skip to content
  1. arXiv Game AI56

    MultiWorldBench: can independently controlled views describe one shared world

    The authors propose MultiWorldBench, a diagnostic Minecraft benchmark for testing whether independently controlled views in a multi-player world model stay consistent with one single persistent shared world; it contains 495 case configurations, seven task suites and ten capabilities.

    Why it matters: With 495 configurations, the paper compares three generative world models against the reference Engine GT on ten capabilities, and the score gaps show where the current weaknesses lie.

  2. arXiv Game AI45

    AgentGarten: a code world framework for continually evolving agents

    The paper proposes the AgentGarten framework, which couples simulators and game engines to a shared neural renderer to build real-time interactive virtual environments. The simulation backend maintains persistent world state and executes program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a unified interface; the neural renderer is adapted from a pretrained video model to take geometry-conditioned input, is distilled with the proposed Adversarial Forcing, and has its inference optimized for real-time interaction.

    Why it matters: The paper wires simulators and game engines into a single neural renderer, and reports learning-efficiency results comparing an agent's 4 rounds of experience with millions of rounds of reinforcement learning.

  1. arXiv Game AI45

    Kuration SDK: Addressing the Virtual2Real Gap in World Models via Data Curation

    The paper proposes and open-sources along with it Kuration SDK, a general-purpose physical AI data curation toolkit that uses data curation to address the Virtual2Real gap in training action-conditioned world models. The authors train and evaluate diffusion world models on CounterStrike gameplay data, confirming that metrics such as FVD, LPIPS and JEDi do not correspond to qualitative playability, and argue that curating raw gameplay data before training begins and measuring multiple diagnostic properties is a more reliable signal.

    Why it matters: The paper trains diffusion world models on CounterStrike gameplay data, points out that visual similarity metrics such as FVD and LPIPS do not correspond to playability, and puts forward a data curation approach.

You've reached the end