--- license: other tags: - counter-strike - world-model - v-jepa2 - video library_name: pytorch --- # CS2 Match-Level Hierarchical World Model (scale300) A causal, action-conditioned latent world model over frozen V-JEPA2 embeddings of **single-POV Counter-Strike 2 gameplay**. Sequences span a player's whole match (rounds concatenated, round-number embedding per token), so the belief state can carry cross-round context. Trained on 300 matches of [RekaAI/CS2-10k](https://huggingface.co/datasets/RekaAI/CS2-10k) (mirage + dust2), 6.94M one-second windows at 0.5s stride. Screen-only contract: the encoder sees POV pixels with no HUD, radar, or parsed game state. Positions and actions are used as training targets and for evaluation, never as inference inputs. ## Files | file | what it is | |---|---| | `hwm_ckpt_scale300.pt` | the world model (12.25M params) + latent norm stats + action norm | | `demo_heads.pt` | readout heads on the same latents: position, position@+8s, behaviour cloning (9 keys), 30s zone intent, zone-transition prior | | `reports/*.json` | every number below, as produced by the scripts | ## Results (match-disjoint holdout, 450 sequences) Latent MSE against the true future latent, versus a persistence ("nothing changes") baseline: | horizon | model | persistence | |---|---|---| | 0.5 s | 0.279 | 0.417 | | 1 s | 0.436 | 0.765 | | 2 s | 0.617 | 1.091 | | 4 s | 0.765 | 1.369 | | 8 s | 0.882 | 1.608 | | 30 s | 0.976 | 1.893 | ## The limitation that matters most Shuffling the action input and re-measuring gives the share of the model's headroom over persistence that actually comes from knowing the action: | horizon | 0.5 s | 1 s | 2 s | 4 s | 8 s | 30 s | |---|---|---|---|---|---|---| | action share of headroom | **24.4%** | 11.6% | 6.6% | 3.1% | 1.2% | **−0.0%** | At 30 s the model is **provably action-blind**, and its MSE of 0.976 on ~unit-variance latents means it is close to a mean predictor there. It beats persistence at 30 s only because persistence is a poor baseline at that range. Structurally: `forward(z, a, rnd)` is causally masked, so token *t* sees actions only up to *t* and each horizon head predicts z*t+h* directly — **there is no slot for a candidate future action sequence**. Counterfactual "what if I do X" requires autoregressive rollout through the h=1 head and is only meaningful for ~0.5–2 s. Macro intent must come from the destination/intent heads, not rollouts. Reproduce with `scripts/hwm/hwm_action_sensitivity.py`. ## 30 s intent probe | features | top-1 | top-3 | moved top-1 | |---|---|---|---| | latent | 0.420 | 0.781 | 0.269 | | HWM context | 0.413 | 0.776 | 0.303 | | latent + context | 0.407 | 0.771 | 0.299 | | *persistence baseline* | *0.371* | — | — | | *zone-transition prior (moved subset)* | — | — | *0.318* | Latent top-1 clears the persistence baseline. On the **moved** subset (63% of cases, where the player actually changes zone — the ones that matter for next-action) all three feature sets still **lose to a trivial zone-transition prior**. `beats_transition_prior` is false everywhere. Scaling 100 → 300 matches narrowed that gap but did not close it, and improved the dynamics model only ~2%, so scaling in this range is sublinear. ## Demo: verified on holdout video `scripts/hwm/demo_pov_video.py` replays a holdout match as real video through the frozen encoder and this checkpoint, with the action input supplied by the BC head reading the same pixels. Ground truth is overlaid only to check. | metric | result | baseline | |---|---|---| | self-localization (median) | 1.99 m | — | | 8 s position forecast | 6.03 m | 9.68 m stand-still | | 30 s intent top-3 | 60.3% | 55.8% zone-transition prior | | 2 s counterfactual, true action closest | 46.2% | 16.7% chance | Caveats, in the interest of not overstating: the 8 s figure shown is a **direct** forecast head — decoding the HWM's own h=16 latent scores slightly worse (6.53 m), so the world model is not what drives that panel. The counterfactual is scored in **latent space**; scored on map positions it falls to 18.6%, i.e. chance, because the position head's ~99 u error swamps the 238 u candidate spread. It is an aggregate claim, not a per-frame one. The position head is unconstrained and emits off-map coordinates on ~1% of frames. ## Intended use Research on latent world models and intent prediction from egocentric video. Not a cheat, not an aim assist: it consumes recorded POV video offline and predicts where a player will be, at metre-to-zone granularity. ## Provenance Derived from RekaAI/CS2-10k; that dataset's terms apply upstream. Encoder is `facebook/vjepa2-vitl-fpc64-256`, frozen. Match split is deterministic (md5 of `match_id`), so the holdout is reproducible. The checkpoint itself is **not** bit-reproducible: training seeds the RNGs but GPU nondeterminism means a re-run yields a statistically equivalent, not identical, model. Code: https://github.com/cristibctr/HWM-CS2