File size: 6,589 Bytes
242cc21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
11ed2c3
 
 
 
 
 
 
 
242cc21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a25356e
242cc21
 
 
 
 
 
2d9ebd2
242cc21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d971c53
 
 
 
 
 
 
 
a5eff8e
d971c53
 
a5eff8e
 
d971c53
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
---
license: cc-by-nc-4.0
tags:
  - world-model
  - video-generation
  - diffusion
  - flow-matching
  - counter-strike
  - reinforcement-learning-environment
library_name: pytorch
---

# CS2 Latent World Model (MIRA-style)

An **action-conditioned world model for Counter-Strike 2** β€” feed it keyboard + mouse actions and it
imagines the next frames. Built by adapting [MIRA](https://github.com/mira-wm/mira) to stream
[RekaAI/CS2-10k](https://huggingface.co/datasets/RekaAI/CS2-10k) (63 TB) with **no local dataset copy**.

<p align="center">
  <img src="assets/walkthrough.gif" alt="Walking around inside the world model" width="640"/>
</p>

<a href="https://hfviewer.com/Quazim0t0/CS2Dream?utm_source=huggingface&amp;utm_medium=embedded_model_card&amp;utm_campaign=Quazim0t0__CS2Dream_card&amp;utm_content=embedded_card_open_viewer&amp;from=embedded-model-card" target="_blank" rel="noopener">
  <img
    src="https://hfviewer.com/api/card.svg?source=Quazim0t0%2FCS2Dream&amp;granularity=auto&amp;v=20260516-title-pills-card"
    alt="Architecture graph for Quazim0t0/CS2Dream. Open in hfviewer"
    width="100%"
  />
</a>

*Driving the world model with "hold W + sweep the mouse" β€” every frame above is **generated**, not real gameplay.*

> **Status: research prototype.** It works end-to-end (coherent, action-responsive rollouts) but is
> small, blurry, and trained on very little data. Limitations are documented honestly below.

## Real vs. generated

| Real gameplay (top) vs. world model (bottom) |
|---|
| ![real vs generated](assets/real_vs_generated_final.png) |

| Codec reconstruction β€” ground truth β€– reconstruction |
|---|
| ![codec reconstruction](assets/codec_reconstruction.png) |

## Architecture

**Stage 1 β€” Video codec (RAEv2)**
- Frozen **DINOv3-L/16** (`facebook/dinov3-vitl16-pretrain-lvd1689m`) backbone, layer aggregation
  over blocks `[11,13,15,17,19,21,23]`.
- Strided-conv bottleneck β†’ **32-channel latent at /32 spatial, /2 temporal**.
- ViT decoder (width 512, depth 6, patch 16). ~331 M params, mostly frozen.

**Stage 2 β€” Latent world model**
- Causal **flow-matching diffusion transformer** over codec latents β€” **12.5 M trainable** params
  (hidden 384, 4 layers, 6 heads).
- AdaLN conditioning on **per-frame keyboard + mouse** actions; diffusion forcing; action dropout 0.1.
- Autoregressive rollout with a streaming KV-cache.

**IO:** 288Γ—512 @ 24 fps, 16-frame clips (CS2 is 48 fps β†’ integer /2 downsample).

## Action space

10 keys β€” `W A S D J C R V [ ]` (move, jump, crouch, walk, freefall, fire, scope) as a multi-hot
vector, **plus continuous mouse deltas**.

> Upstream MIRA conditions on *key presses only* (Rocket League is keyboard-only). CS2 is
> mouse-driven, so mouse deltas are populated here and consumed by MIRA's existing `ActionEncoder`
> (symlog-normalized + temporally pooled).

## Files

| File | What |
|---|---|
| `codec_069000.pt` | Codec weights (decoder + bottleneck; frozen DINOv3 rebuilt from HF) |
| `world_model_030000.pt` | World model weights (12.5 M) |
| `scripts/`, `src/`, `configs/` | Training / eval / streaming code |
| `patches/dino_hf_backbone.patch` | Load gated DINOv3 via `transformers` instead of Meta's `.pth` |


## Training

| | Codec | World model |
|---|---|---|
| Steps | 69,000 | 30,000 |
| Hardware | RTX 5060 Ti (16 GB) | RTX 5060 Ti |
| Loss | L1 reconstruction | flow-matching diffusion |
| Final | ema 0.1392 | ema 2.41 |

Trained on the untarred `sample/` split: **3 matches / 374 clips / 44 (match, round) groups.**

## Results

- **Codec reconstruction: 24.75 dB PSNR** (vs 19.7 dB with a random backbone). Recovers room
  geometry and recognizable player figures; no fine texture.
- **Rollouts stay coherent** β€” players persist across the clip, room geometry is stable.
- **Controllability** β€” mean abs pixel difference of a generated rollout vs an idle-action rollout:
  - `forward` vs `idle`: **0.025 – 0.037**
  - `turn` vs `idle`: **0.040**
  - Non-zero **and growing over the clip** β†’ the model integrates actions over time rather than
    ignoring them. This was the go/no-go test, and it passed.

## Limitations

- **Blurry.** The /32, 32-channel latent + a depth-6 decoder caps detail (no faces, weapon detail,
  floor texture). The paper uses a ViT-XL decoder.
- **Long-horizon drift.** Coherent for several frames past context, then degrades β€” dissolving to
  generic wall texture or darkening. 16-frame clips vs the paper's 80.
- **Tiny dataset β€” 3 matches.** The dominant constraint; there's little to generalize from.
- **Tiny model** β€” 12.5 M vs the paper's 1 B. Long-horizon coherence is what capacity buys.
- **Single-perspective only.** MIRA's defining contribution is *multiplayer* (joint, time-aligned
  multi-POV with per-player action attribution). Only the single-player stage is trained here. The
  multi-POV data path exists (`cs2_stream` groups POVs by `(match_id, round_number)`, with ranged
  tar-member fetch for the full split) but the joint stage was never trained.
- **Action semantics unverified.** We measure that actions *change* the output, not that "forward"
  moves precisely forward β€” fidelity is too low to confirm direction visually.

## Intended use / out of scope

Research on world models and action-conditioned video prediction. **Not** a game, an anti-cheat
signal, a player model, or anything deployable. Do not present generated frames as real gameplay.

## Attribution & licensing

- **MIRA** β€” architecture/training code (upstream; this repo ships patches + new modules).
- **DINOv3** (Meta) β€” frozen backbone, **gated weights** under Meta's license (not redistributed here).
- **CS2-10k** (RekaAI) β€” **CC BY-NC 4.0**.
- Counter-Strike 2 imagery Β© Valve; generated frames derive from that footage.

**Non-commercial research use only.**

## Usage

```bash
export RS_DINO_HF=facebook/dinov3-vitl16-pretrain-lvd1689m   # requires HF access to gated DINOv3
python scripts/consolidate_codec.py     # pack codec_069000.pt for the world model
python scripts/train_cs2_wm.py --test   # controllability check (generates rollouts)
python scripts/eval_cs2_codec.py        # reconstruction PSNR + side-by-side frames
```

<!-- CITE_START -->
## Citation

If you use this model, please cite:

```bibtex
@misc{escarda86mbase,
  title        = {CS2Dream: Counterstrike small latent world model},
  author       = {Dean Byrne (Quazim0t0)},
  year         = {2026},
  howpublished = {HuggingFace, \url{https://huggingface.co/Quazim0t0/CS2Dream}},
  note         = {Quazim0t0/CS2Dream}
}
```
<!-- CITE_END -->