File size: 6,589 Bytes
242cc21 11ed2c3 242cc21 a25356e 242cc21 2d9ebd2 242cc21 d971c53 a5eff8e d971c53 a5eff8e d971c53 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 | ---
license: cc-by-nc-4.0
tags:
- world-model
- video-generation
- diffusion
- flow-matching
- counter-strike
- reinforcement-learning-environment
library_name: pytorch
---
# CS2 Latent World Model (MIRA-style)
An **action-conditioned world model for Counter-Strike 2** β feed it keyboard + mouse actions and it
imagines the next frames. Built by adapting [MIRA](https://github.com/mira-wm/mira) to stream
[RekaAI/CS2-10k](https://huggingface.co/datasets/RekaAI/CS2-10k) (63 TB) with **no local dataset copy**.
<p align="center">
<img src="assets/walkthrough.gif" alt="Walking around inside the world model" width="640"/>
</p>
<a href="https://hfviewer.com/Quazim0t0/CS2Dream?utm_source=huggingface&utm_medium=embedded_model_card&utm_campaign=Quazim0t0__CS2Dream_card&utm_content=embedded_card_open_viewer&from=embedded-model-card" target="_blank" rel="noopener">
<img
src="https://hfviewer.com/api/card.svg?source=Quazim0t0%2FCS2Dream&granularity=auto&v=20260516-title-pills-card"
alt="Architecture graph for Quazim0t0/CS2Dream. Open in hfviewer"
width="100%"
/>
</a>
*Driving the world model with "hold W + sweep the mouse" β every frame above is **generated**, not real gameplay.*
> **Status: research prototype.** It works end-to-end (coherent, action-responsive rollouts) but is
> small, blurry, and trained on very little data. Limitations are documented honestly below.
## Real vs. generated
| Real gameplay (top) vs. world model (bottom) |
|---|
|  |
| Codec reconstruction β ground truth β reconstruction |
|---|
|  |
## Architecture
**Stage 1 β Video codec (RAEv2)**
- Frozen **DINOv3-L/16** (`facebook/dinov3-vitl16-pretrain-lvd1689m`) backbone, layer aggregation
over blocks `[11,13,15,17,19,21,23]`.
- Strided-conv bottleneck β **32-channel latent at /32 spatial, /2 temporal**.
- ViT decoder (width 512, depth 6, patch 16). ~331 M params, mostly frozen.
**Stage 2 β Latent world model**
- Causal **flow-matching diffusion transformer** over codec latents β **12.5 M trainable** params
(hidden 384, 4 layers, 6 heads).
- AdaLN conditioning on **per-frame keyboard + mouse** actions; diffusion forcing; action dropout 0.1.
- Autoregressive rollout with a streaming KV-cache.
**IO:** 288Γ512 @ 24 fps, 16-frame clips (CS2 is 48 fps β integer /2 downsample).
## Action space
10 keys β `W A S D J C R V [ ]` (move, jump, crouch, walk, freefall, fire, scope) as a multi-hot
vector, **plus continuous mouse deltas**.
> Upstream MIRA conditions on *key presses only* (Rocket League is keyboard-only). CS2 is
> mouse-driven, so mouse deltas are populated here and consumed by MIRA's existing `ActionEncoder`
> (symlog-normalized + temporally pooled).
## Files
| File | What |
|---|---|
| `codec_069000.pt` | Codec weights (decoder + bottleneck; frozen DINOv3 rebuilt from HF) |
| `world_model_030000.pt` | World model weights (12.5 M) |
| `scripts/`, `src/`, `configs/` | Training / eval / streaming code |
| `patches/dino_hf_backbone.patch` | Load gated DINOv3 via `transformers` instead of Meta's `.pth` |
## Training
| | Codec | World model |
|---|---|---|
| Steps | 69,000 | 30,000 |
| Hardware | RTX 5060 Ti (16 GB) | RTX 5060 Ti |
| Loss | L1 reconstruction | flow-matching diffusion |
| Final | ema 0.1392 | ema 2.41 |
Trained on the untarred `sample/` split: **3 matches / 374 clips / 44 (match, round) groups.**
## Results
- **Codec reconstruction: 24.75 dB PSNR** (vs 19.7 dB with a random backbone). Recovers room
geometry and recognizable player figures; no fine texture.
- **Rollouts stay coherent** β players persist across the clip, room geometry is stable.
- **Controllability** β mean abs pixel difference of a generated rollout vs an idle-action rollout:
- `forward` vs `idle`: **0.025 β 0.037**
- `turn` vs `idle`: **0.040**
- Non-zero **and growing over the clip** β the model integrates actions over time rather than
ignoring them. This was the go/no-go test, and it passed.
## Limitations
- **Blurry.** The /32, 32-channel latent + a depth-6 decoder caps detail (no faces, weapon detail,
floor texture). The paper uses a ViT-XL decoder.
- **Long-horizon drift.** Coherent for several frames past context, then degrades β dissolving to
generic wall texture or darkening. 16-frame clips vs the paper's 80.
- **Tiny dataset β 3 matches.** The dominant constraint; there's little to generalize from.
- **Tiny model** β 12.5 M vs the paper's 1 B. Long-horizon coherence is what capacity buys.
- **Single-perspective only.** MIRA's defining contribution is *multiplayer* (joint, time-aligned
multi-POV with per-player action attribution). Only the single-player stage is trained here. The
multi-POV data path exists (`cs2_stream` groups POVs by `(match_id, round_number)`, with ranged
tar-member fetch for the full split) but the joint stage was never trained.
- **Action semantics unverified.** We measure that actions *change* the output, not that "forward"
moves precisely forward β fidelity is too low to confirm direction visually.
## Intended use / out of scope
Research on world models and action-conditioned video prediction. **Not** a game, an anti-cheat
signal, a player model, or anything deployable. Do not present generated frames as real gameplay.
## Attribution & licensing
- **MIRA** β architecture/training code (upstream; this repo ships patches + new modules).
- **DINOv3** (Meta) β frozen backbone, **gated weights** under Meta's license (not redistributed here).
- **CS2-10k** (RekaAI) β **CC BY-NC 4.0**.
- Counter-Strike 2 imagery Β© Valve; generated frames derive from that footage.
**Non-commercial research use only.**
## Usage
```bash
export RS_DINO_HF=facebook/dinov3-vitl16-pretrain-lvd1689m # requires HF access to gated DINOv3
python scripts/consolidate_codec.py # pack codec_069000.pt for the world model
python scripts/train_cs2_wm.py --test # controllability check (generates rollouts)
python scripts/eval_cs2_codec.py # reconstruction PSNR + side-by-side frames
```
<!-- CITE_START -->
## Citation
If you use this model, please cite:
```bibtex
@misc{escarda86mbase,
title = {CS2Dream: Counterstrike small latent world model},
author = {Dean Byrne (Quazim0t0)},
year = {2026},
howpublished = {HuggingFace, \url{https://huggingface.co/Quazim0t0/CS2Dream}},
note = {Quazim0t0/CS2Dream}
}
```
<!-- CITE_END --> |