VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
Key results:
- π Covers all four stages that decide planning success: encoder, predictor, planner, task tolerance
- π₯ High correlation with success rate: pooled Ο +0.90, against +0.82 for raw empowerment
- πͺ Best calibration across planners and datasets: below the constant-predictor reference on every pool
- π 20Γ faster than the real evaluation: ~7 s per checkpoint, no environment and no planner rollouts
- π§© Model-agnostic: three methods (
encode,action_embed,predict_next) are the whole interface - β»οΈ CPU-reproducible: every table in the paper recomputes in seconds
Available Checkpoints
Each file is a full-object torch.save pickle of jepa.JEPA: ViT-Tiny encoder + action-conditioned AdaLN transformer predictor, ~15M parameters, 72 MB. Filenames say which model the weights are β vis-wm_epoch_<N>.ckpt, lewm_epoch_<N>.ckpt (the SIGReg baseline) or dino-cls_epoch_<N>.ckpt.
| Path | Contents | Files | Size |
|---|---|---|---|
vis-wm/<task>/seed<S>/ |
VIS-WM at the epoch the paper reports, + config.yaml |
39 | 0.88 GiB |
baselines-lewm/maze2d/seed729/ |
SIGReg (LeWM) baseline on MAZE | 3 | 0.07 GiB |
pools/<pool>/<env>/<run>/ |
every checkpoint behind a reported number, plus the calibration pool | 547 | 36.7 GiB |
VIS-WM: PushT / Reacher / Two-Room / Cube Γ seeds 403, 46, 729, plus MAZE at seed 729. Reported epoch per arm β maximum success on the single-seed d=25 evaluation, ties to the later epoch:
| Task | seed 403 | seed 46 | seed 729 |
|---|---|---|---|
| PushT | 9 | 10 | 10 |
| Reacher | 5 | 10 | 4 |
| Two-Room | 5 | 10 | 9 |
| Cube | 6 | 8 | 4 |
| MAZE | β | β | 8 (SIGReg baseline: 10) |
Pools. The split is by training run, so no checkpoint in a test pool comes from a run used to fit anything. Every directory says what its checkpoints are for:
| Directory | Ckpt / runs | Which reported number it produces | Evaluation seeds |
|---|---|---|---|
pools/development/ |
137 / 14 | block 1 of tab:viscore; the constants are fitted here |
all available |
pools/heldout/ |
103 / 33 | block 2 -- the headline test, runs disjoint from development | 5501, 60601, 90210 |
pools/heldout-method/ |
3 / 3 | block 3, our DINO-CLS arms (the other 27 are other groups', not re-hosted) | its own |
pools/heldout-dataset/ |
20 / 2 | block 4 -- the unseen MAZE task | its own |
pools/cube-reference/ |
50 / 5 | the parenthesised Cube column, and nothing else | 5501, 60601, 90210 |
pools/calibration-fit/ |
232 / 47 | no reported number: these fit the frozen map blocks 3--4 apply | all available |
So: 340 checkpoints produce every number in the paper (313 hosted here + 27 fetched with
reproduce/download_external.py), 572 if you also want to refit the frozen calibration map
from weights, and 0 if you only want to recompute the tables from the shipped metric values.
Per-task counts in the two ranked pools (PushT / Reacher / Two-Room): development 57 / 50 / 30, held-out 27 / 44 / 32.
heldout absorbed what earlier versions reported as a separate terminal pool: 14 of that pool's
17 runs were also in the held-out fold, and restricted to converged epochs its labels no longer
resolve individual checkpoints.
cube-reference is separate from heldout because Cube entered no fold. Its label spread does not
exceed its own binomial standard error, so nothing there is rankable; the paper prints its
correlation in parentheses and excludes it from every pooled and calibration number. These
checkpoints are published because that exclusion argument is measured on them.
calibration-fit is the part of the 472-cell calibration pool that is not already in
development or heldout (both of which are subsets of it). Its checkpoints enter no correlation
and no reported error -- they only determine the metric-to-success map that the two transfer blocks
apply without refitting. They are published so that a competing metric can be calibrated on the
same 472 cells.
Three SIGReg Two-Room runs landed in neither fold and are not released, which is why Two-Room
carries no SIGReg checkpoint in heldout.
The rest of the held-out-method pool is other groups' released checkpoints (Qantara, RC-aux,
INTACT), which are not re-hosted here; fetch them with reproduce/download_external.py.
pools/pool_manifest.csv gives each checkpoint's pool membership, success labels and seven metric values; pools/pool_assignment.csv gives run β fold.
Training
LeWM architecture with the VISReg regularizer, Ξ» = 4.5. AdamW, linear warmup + cosine decay, 10 epochs, batch 128, lr 1e-4. Two-Room uses history_size=1, the other tasks 3.
Usage
Score a checkpoint
from viscore import score_checkpoint
f = score_checkpoint("vis-wm/pusht/seed403/vis-wm_epoch_9.ckpt",
"probes/probe_pusht.npz", task="pusht")
print(f) # VIS + veracity / influence / sobriety + their inputs
The pickle names jepa and module as top-level modules; viscore binds the shipped copies automatically, so no LeWorldModel checkout is needed.
Download with huggingface_hub
from huggingface_hub import hf_hub_download, snapshot_download
# one checkpoint
path = hf_hub_download(repo_id="BooBooWu/viscore",
filename="vis-wm/pusht/seed403/vis-wm_epoch_9.ckpt")
# a whole tree
snapshot_download(repo_id="BooBooWu/viscore", allow_patterns=["vis-wm/*"])
Build a probe
A probe is one frozen slice of training data (rng(0), 300 episodes, frameskip 5) shared by every checkpoint compared.
viscore probe --tasks pusht --data-home $STABLEWM_HOME --out-dir probes
viscore score --task pusht --probe probes/probe_pusht.npz \
--run-dir vis-wm/pusht/seed403 --epochs 9 --csv out.csv
Evaluation
Both of the paper's main tables recompute from files shipped in the GitHub repo β no GPU, dataset or checkpoint download:
git clone https://github.com/HaiyuWu/viscore && cd viscore && pip install -e .
python reproduce/tables.py # metric vs success, three pools
python reproduce/planning_tables.py --strict # planning tables, gated against published values
Notes: peak epoch = maximum evaluation success rate, ties to the later epoch. Cube is reported in parentheses and excluded from pooled and calibration columns. All published labels and sobriety probes use CEM. A single 50-episode evaluation has a success sd of β6 points.
Citation
@article{wu2026viscore,
title = {VIScore: Diagnosing Planning-Relevant Quality in Latent World Models},
author = {Wu, Haiyu and Balestriero, Randall and Levine, Morgan},
journal = {arXiv preprint arXiv:2608.11174},
year = {2026},
eprint = {2608.11174},
archivePrefix = {arXiv}
}
Architecture and training recipe from LeWorldModel; regularizer from VISReg.
License
This project (code and checkpoints) is released under the MIT License.