Skip to content
TodayOct 9Fri2 items
  1. arXiv Game AI49

    Can AI agents learn their way to the top? AAArena evaluates heuristic learning in a long-running game competition

    The paper proposes the AAArena benchmark, using 12 real adversarial games and 1,920 archived human programs to evaluate agents' learning ability in long-running competitions. Among the evaluated model and tool configurations, Opus5.5 with Claude Code took 6 gold medals, and no configuration could take the remaining 6 human ladders.

    Why it matters: AAArena organizes 12 adversarial games and 1,920 human programs into a competition-style evaluation, which can be used to observe the limits of how agents improve their strategies from limited samples.

  2. arXiv Game AI45

    AgentGarten: a code world framework for continually evolving agents

    The paper proposes the AgentGarten framework, which couples simulators and game engines to a shared neural renderer to build real-time interactive virtual environments. The simulation backend maintains persistent world state and executes program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a unified interface; the neural renderer is adapted from a pretrained video model to take geometry-conditioned input, is distilled with the proposed Adversarial Forcing, and has its inference optimized for real-time interaction.

    Why it matters: The paper wires simulators and game engines into a single neural renderer, and reports learning-efficiency results comparing an agent's 4 rounds of experience with millions of rounds of reinforcement learning.

Oct 7Wed
  1. arXiv Game AI45

    Kuration SDK: Addressing the Virtual2Real Gap in World Models via Data Curation

    The paper proposes and open-sources along with it Kuration SDK, a general-purpose physical AI data curation toolkit that uses data curation to address the Virtual2Real gap in training action-conditioned world models. The authors train and evaluate diffusion world models on CounterStrike gameplay data, confirming that metrics such as FVD, LPIPS and JEDi do not correspond to qualitative playability, and argue that curating raw gameplay data before training begins and measuring multiple diagnostic properties is a more reliable signal.

    Why it matters: The paper trains diffusion world models on CounterStrike gameplay data, points out that visual similarity metrics such as FVD and LPIPS do not correspond to playability, and puts forward a data curation approach.

  2. arXiv Game AI29

    LeCuration: A Tiny World Model as a Data Curation Multi-Tool

    LeCuration is a tiny world model used to curate data for another, larger downstream model. It uses LeWorldModel (LeWM) as a latent encoder and predictor, and adds a DiT decoder to supply visuals for autoregressive gameplay rollouts; its embeddings can serve as an anomaly detection signal and a content-based clustering heuristic, while autoregressive prediction of game state can be used to qualitatively check whether actions and states are consistent. The team ran a qualitative proof of concept on CS:GO gameplay data and has not yet reported quantitative curation metrics or downstream training results.

  3. arXiv Game AI47

    Paper proposes Recursive Game Creator, using four recursively iterating components to improve generated game experience

    The paper proposes Recursive Game Creator, an experience-oriented agentic game development framework that uses four components — Designer, Builder, Player and Reviewer — to iterate recursively and push a rough game prototype toward a more replayable work.

    Why it matters: The paper presents the four-component recursive workflow and results on two benchmarks, with 53.2% success rate on strict tasks in GameASG-Bench, a 34.1% improvement over the same-model baseline.

  4. arXiv Game AI27

    A space-agnostic visual game analytics tool for HoloLens 2: adaptive spatial reconstruction and realtime object detection

    A space-agnostic visual game analytics tool for mixed reality (MR) game development has been proposed, combining adaptive spatial reconstruction with realtime object detection and running on the Microsoft HoloLens 2 MR headset. The tool is designed for the complexity of MR games, which must integrate real-world and digital components, and for the difficulty existing visual analytics tools have in achieving space-agnostic generalization. The researchers hope it will deepen understanding of players' surroundings and open a new direction for visual analytics in MR game development.

  5. arXiv Game AI36

    Study proposes ResNet-BiLSTM for multi-label perceptual bug detection from gameplay footage

    The study proposes a ResNet-BiLSTM model that performs multi-label perceptual bug detection on gameplay footage, achieving an F1 score of 85.78% on a benchmark dataset, and compares it with video classification models including Inflated 3D ConvNet and 3D ResNet.

    Why it matters: The paper reports an F1 of 85.78% for ResNet-BiLSTM and compares it with video classification models such as Inflated 3D ConvNet on the same task.

Oct 6Tue
  1. arXiv Game AI42

    SpeedrunBench launches to benchmark frontier LLM agents' strategy formation with speedruns across 9 games

    The paper introduces SPEEDRUNBENCH, a speedrunning benchmark that evaluates frontier LLM agents on 9 different games. Agents must repeatedly refine their strategies over long-horizon actions, review their own performance and use the knowledge they have gained to be faster than their past selves and others. Experiments show frontier agents approach human world records on simple platformers, but on longer, more complex games they still lag behind human performance under practical budget constraints; the authors see the benchmark as a long-term testbed for studying agents' strategy formation ability.

    Why it matters: The paper uses speedrun tasks across 9 games to probe agents' strategy formation, showing near-human world-record performance on simple platformers while still trailing on long, complex games.

  2. arXiv Game AI41

    Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

    The researchers propose Attacca, which trains a vision goal-conditioned policy on complete search-to-interaction trajectories so long-horizon embodied agents can carry on to the next task from the position, orientation and world state left by the previous one. The method uses context-decoupled goal sampling, pairing each demonstration with a class-compatible masked goal image from another world, and provides supervision beyond action imitation through a goal mask prediction head, while introducing behavior phase conditioning that distinguishes the Search, Approach and Interact phases.

    Why it matters: The paper brings state-continuous long-horizon execution into the evaluation and reports success rate and completion comparisons against the strongest baseline.

  3. arXiv Game AI31

    SENSE: a state-aware emotion navigation storytelling engine

    The paper proposes SENSE, a state-aware framework for generating playable branching visual novels with multi-track affective navigation. The framework integrates a stateful narrative architecture called MIND, a structural analyzer and a path-aware context management module, and can generate multiple intersecting routes from a small number of high-level inputs while maintaining character consistency and narrative causality. Evaluations using LLM judges, affective metrics and visual evaluation show SENSE outperforms baselines in narrative diversity and asset integration robustness, and initial human playtests show a directional improvement in affective fidelity with comparable fun.

  4. arXiv Game AI35

    Emoception proposes the SALFT framework, selectively fine-tuning a video ViT to recognize changes in player arousal

    The paper proposes SALFT (Selective Affective Layer Fine-Tuning), a framework that fine-tunes a Video Vision Transformer using a selection criterion based on layer-wise parameter L2 norm change to recognize changes in player arousal from gameplay footage.

    Why it matters: The paper uses layer-wise parameter L2 norm change as its selection criterion and shows that updating only about 8% of parameters approaches full fine-tuning, useful for comparing compute costs against your own player emotion recognition pipeline.

Oct 5Mon