doom-world-model

A diffusion world model of Doom, trained from scratch on the first level of Freedoom (E1M1) in ViZDoom. Given the last 4 frames and actions, a small U-Net denoises the next 128×96 frame; fed its own frames, it generates the level as you walk, turn and shoot through it. There is no game engine at play time, only the model. After DIAMOND and GameNGen, at a much smaller scale. doom.py in this repo collects the data, trains, plays and doubles as the loader.

Real frames (left) and the model's rollout from the same start and actions (right)

Left: real game. Right: the model, from the same four starting frames and the same actions, never seeing a real frame again.

🚀 Usage

Play it in the browser. Each button press streams four generated frames, about half a second of game time:

uv run "$(hf download jgalego/doom-world-model doom.py --quiet)" play

From Python, with doom.py on the path:

import torch
from doom import DEVICE, collect, load, rollout, to_float

model = load("jgalego/doom-world-model")
size = (model.config.width, model.config.height)
frames, _, _ = collect(4, seed=0, size=size)  # four real frames to start from
context = to_float(frames[:4])[None].to(DEVICE)
moves = torch.tensor([[0, 0, 0] + [1] * 10 + [3] * 6], device=DEVICE)  # FORWARD, then LEFT
video = rollout(model, context, moves)  # (1, 16, 3, height, width) in [-1, 1]

Actions: 0 NOOP, 1 FORWARD, 2 BACKWARD, 3 LEFT (turn), 4 RIGHT (turn), 5 ATTACK, 6 USE (doors, switches). One step is 4 game tics, about 0.11 s.

🏋️ Training

Model U-Net, channels 64, 128, 256, 256, 7.5M parameters; noise level and the last 4 actions modulate every block
Diffusion EDM preconditioning, σ_data 0.5, log-normal training noise (log σ mean 0.4, std 1.2); 3 Euler steps per frame from σ=5.0
Data 200,000 frames (217 episodes of up to two minutes) of a random policy on Freedoom E1M1: it holds each action for 1 to 6 steps, mostly walks forward, and turns away or tries USE when the depth buffer shows a wall ahead
Steps 5,000, batch size 64
Optimizer AdamW, lr 0.0001, 1000 warmup steps, cosine decay, gradient clipping 1.0; weights are an EMA with decay 0.999
Hardware NVIDIA A10G, 16.5 min

📊 Results

Rollouts on 256 windows from episodes whose seeds never appear in training; they play the same level. The model gets four real frames and the real actions, then only its own frames. Repeat keeps showing the last real frame. Changed pixels scores only the pixels where the real or the predicted frame differs from the last real frame, so the HUD and still views do not dominate.

Frames ahead Model, all pixels (dB) Repeat, all pixels (dB) Model, changed pixels (dB) Repeat, changed pixels (dB)
1 22.36 21.29 19.77 16.2
5 17.55 19.2 16.62 15.5
15 14.01 18.16 13.53 15.09

⚠️ Limitations

  • Rollouts collapse. From four real frames, the first generated frame is already a brown haze and by the fifth only a grey wall, the gun and a breaking HUD remain (see the GIF). It still beats repeating the last frame on PSNR up to 5 frames ahead, because a grey average scores well on PSNR. The model learned to denoise frames but not to predict the next one from context.
  • Training longer makes it worse. Held-out rollout quality peaks after about 4,000 steps and then declines while the training loss keeps falling. A 50,000-step run, an EMA of the weights, fp32 instead of bf16, warmup with cosine decay and 64×48 frames did not change that.
  • One level, seen through random play: the model knows the parts of E1M1 a random walker reaches, not the whole map, and nothing of other levels.
  • Four frames of memory, under half a second: walk out of a room and back and it may not be the same room.
  • 128×96 frames, three denoising steps: textures blur and enemies smear, more so the longer it runs.
  • Errors compound: every generated frame becomes input for the next.

Freedoom assets are under the BSD 3-clause license.

Downloads last month
-
Safetensors
Model size
7.53M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including jgalego/doom-world-model

Papers for jgalego/doom-world-model

Evaluation results

  • PSNR on changed pixels at 1 frames on Freedoom E1M1, held-out random play
    self-reported
    19.770
  • PSNR on changed pixels at 5 frames on Freedoom E1M1, held-out random play
    self-reported
    16.620
  • PSNR on changed pixels at 15 frames on Freedoom E1M1, held-out random play
    self-reported
    13.530