doom-world-model
A diffusion world model of Doom, trained from scratch on the first level of Freedoom (E1M1) in ViZDoom. Given the last 4 frames and actions, a small U-Net denoises the next 128×96 frame; fed its own frames, it generates the level as you walk, turn and shoot through it. There is no game engine at play time, only the model. After DIAMOND and GameNGen, at a much smaller scale. doom.py in this repo collects the data, trains, plays and doubles as the loader.
Left: real game. Right: the model, from the same four starting frames and the same actions, never seeing a real frame again.
🚀 Usage
Play it in the browser. Each button press streams four generated frames, about half a second of game time:
uv run "$(hf download jgalego/doom-world-model doom.py --quiet)" play
From Python, with doom.py on the path:
import torch
from doom import DEVICE, collect, load, rollout, to_float
model = load("jgalego/doom-world-model")
size = (model.config.width, model.config.height)
frames, _, _ = collect(4, seed=0, size=size) # four real frames to start from
context = to_float(frames[:4])[None].to(DEVICE)
moves = torch.tensor([[0, 0, 0] + [1] * 10 + [3] * 6], device=DEVICE) # FORWARD, then LEFT
video = rollout(model, context, moves) # (1, 16, 3, height, width) in [-1, 1]
Actions: 0 NOOP, 1 FORWARD, 2 BACKWARD, 3 LEFT (turn), 4 RIGHT (turn), 5 ATTACK, 6 USE (doors, switches). One step is 4 game tics, about 0.11 s.
🏋️ Training
| Model | U-Net, channels 64, 128, 256, 256, 7.5M parameters; noise level and the last 4 actions modulate every block |
| Diffusion | EDM preconditioning, σ_data 0.5, log-normal training noise (log σ mean 0.4, std 1.2); 3 Euler steps per frame from σ=5.0 |
| Data | 200,000 frames (217 episodes of up to two minutes) of a random policy on Freedoom E1M1: it holds each action for 1 to 6 steps, mostly walks forward, and turns away or tries USE when the depth buffer shows a wall ahead |
| Steps | 5,000, batch size 64 |
| Optimizer | AdamW, lr 0.0001, 1000 warmup steps, cosine decay, gradient clipping 1.0; weights are an EMA with decay 0.999 |
| Hardware | NVIDIA A10G, 16.5 min |
📊 Results
Rollouts on 256 windows from episodes whose seeds never appear in training; they play the same level. The model gets four real frames and the real actions, then only its own frames. Repeat keeps showing the last real frame. Changed pixels scores only the pixels where the real or the predicted frame differs from the last real frame, so the HUD and still views do not dominate.
| Frames ahead | Model, all pixels (dB) | Repeat, all pixels (dB) | Model, changed pixels (dB) | Repeat, changed pixels (dB) |
|---|---|---|---|---|
| 1 | 22.36 | 21.29 | 19.77 | 16.2 |
| 5 | 17.55 | 19.2 | 16.62 | 15.5 |
| 15 | 14.01 | 18.16 | 13.53 | 15.09 |
⚠️ Limitations
- Rollouts collapse. From four real frames, the first generated frame is already a brown haze and by the fifth only a grey wall, the gun and a breaking HUD remain (see the GIF). It still beats repeating the last frame on PSNR up to 5 frames ahead, because a grey average scores well on PSNR. The model learned to denoise frames but not to predict the next one from context.
- Training longer makes it worse. Held-out rollout quality peaks after about 4,000 steps and then declines while the training loss keeps falling. A 50,000-step run, an EMA of the weights, fp32 instead of bf16, warmup with cosine decay and 64×48 frames did not change that.
- One level, seen through random play: the model knows the parts of E1M1 a random walker reaches, not the whole map, and nothing of other levels.
- Four frames of memory, under half a second: walk out of a room and back and it may not be the same room.
- 128×96 frames, three denoising steps: textures blur and enemies smear, more so the longer it runs.
- Errors compound: every generated frame becomes input for the next.
Freedoom assets are under the BSD 3-clause license.
- Downloads last month
- -
Collection including jgalego/doom-world-model
Papers for jgalego/doom-world-model
Diffusion Models Are Real-Time Game Engines
Diffusion for World Modeling: Visual Details Matter in Atari
Evaluation results
- PSNR on changed pixels at 1 frames on Freedoom E1M1, held-out random playself-reported19.770
- PSNR on changed pixels at 5 frames on Freedoom E1M1, held-out random playself-reported16.620
- PSNR on changed pixels at 15 frames on Freedoom E1M1, held-out random playself-reported13.530
