Spaces:
Running on Zero
Running on Zero
| title: MMBench2 · Hallucination Signals vs True Error | |
| emoji: 🔭 | |
| colorFrom: indigo | |
| colorTo: gray | |
| sdk: gradio | |
| sdk_version: 6.17.3 | |
| app_file: app.py | |
| python_version: "3.10" | |
| startup_duration_timeout: 40m | |
| pinned: false | |
| license: mit | |
| short_description: Hallucination signals vs. true world-model rollout error | |
| # Hallucination signals vs. true rollout error | |
| Interactive companion to *"Hallucination in World Models is Predictable and | |
| Preventable"*. Drive a 350M-parameter world model with the keyboard and watch four | |
| signals get scored, live, against the error the model actually makes. | |
| | signal | what it measures | sees the future? | | |
| |---|---|---| | |
| | `u_r` | motion-normalized tokenizer round-trip residual | no | | |
| | `u_f` | denoising-trajectory instability | no | | |
| | `u_s` | motion-normalized inter-seed variance | no | | |
| | **`WAV`** | **inverse-dynamics rollout divergence (ours)** | **yes** | | |
| **Target.** Pearson *r* against the mean distance between an open-loop rollout | |
| driven by the *ground-truth* actions and the real trajectory, unnormalized. | |
| The target is deliberately not motion-normalized. `u_r` and `u_s` are themselves | |
| divided by a motion term, so scoring them against a motion-normalized target | |
| shares that 1/motion factor and manufactures correlation. Measured offline over | |
| 120 trajectories, switching the target from normalized to unnormalized moved | |
| `u_r` from 0.715 to 0.394 and `u_s` from 0.818 to 0.490 while `u_f` *rose* from | |
| 0.591 to 0.705 — it reversed the ranking. | |
| **WAV is not a peer of the `u_*` signals** and the UI keeps it apart. It infers | |
| the action from the real next frames, so it is an offline audit measurement, not | |
| a runtime predictor: its ceiling is 1.0, reached by using the true action instead | |
| of the inferred one. Over 150 trajectories on 6 tasks it scores **0.959**, with | |
| the label-free estimate landing within **1%** of the labelled measurement. | |
| ## Controls | |
| | input | action | | |
| |---|---| | |
| | arrows / WASD | actions (dimensions depend on the task) | | |
| | Space, or click the frame | pause | | |
| | R | reset the episode | | |
| | dropdown | switch task | | |
| Correlations need a few rollout samples before they appear — the multi-step | |
| rollout is scored with a lag, since the real frames it is compared against have | |
| to arrive first. | |
| ## Notes | |
| GPU time is allocated in bounded windows. The session survives a window ending, | |
| so the step counter keeps climbing and your accumulated statistics are not reset. | |
| Task domains whose native dependencies fail to build are hidden from the dropdown | |
| rather than breaking the app; the startup log lists any that were dropped. | |
| Configurable via Space variables: `WAV_EVERY`, `WAV_HORIZON`, `WAV_HORIZON_LONG`, | |
| `GPU_DURATION`, `WM_VARIANT`. | |