--- title: MMBench2 ยท Hallucination Signals vs True Error emoji: ๐Ÿ”ญ colorFrom: indigo colorTo: gray sdk: gradio sdk_version: 6.17.3 app_file: app.py python_version: "3.10" startup_duration_timeout: 40m pinned: false license: mit short_description: Hallucination signals vs. true world-model rollout error --- # Hallucination signals vs. true rollout error Interactive companion to *"Hallucination in World Models is Predictable and Preventable"*. Drive a 350M-parameter world model with the keyboard and watch four signals get scored, live, against the error the model actually makes. | signal | what it measures | sees the future? | |---|---|---| | `u_r` | motion-normalized tokenizer round-trip residual | no | | `u_f` | denoising-trajectory instability | no | | `u_s` | motion-normalized inter-seed variance | no | | **`WAV`** | **inverse-dynamics rollout divergence (ours)** | **yes** | **Target.** Pearson *r* against the mean distance between an open-loop rollout driven by the *ground-truth* actions and the real trajectory, unnormalized. The target is deliberately not motion-normalized. `u_r` and `u_s` are themselves divided by a motion term, so scoring them against a motion-normalized target shares that 1/motion factor and manufactures correlation. Measured offline over 120 trajectories, switching the target from normalized to unnormalized moved `u_r` from 0.715 to 0.394 and `u_s` from 0.818 to 0.490 while `u_f` *rose* from 0.591 to 0.705 โ€” it reversed the ranking. **WAV is not a peer of the `u_*` signals** and the UI keeps it apart. It infers the action from the real next frames, so it is an offline audit measurement, not a runtime predictor: its ceiling is 1.0, reached by using the true action instead of the inferred one. Over 150 trajectories on 6 tasks it scores **0.959**, with the label-free estimate landing within **1%** of the labelled measurement. ## Controls | input | action | |---|---| | arrows / WASD | actions (dimensions depend on the task) | | Space, or click the frame | pause | | R | reset the episode | | dropdown | switch task | Correlations need a few rollout samples before they appear โ€” the multi-step rollout is scored with a lag, since the real frames it is compared against have to arrive first. ## Notes GPU time is allocated in bounded windows. The session survives a window ending, so the step counter keeps climbing and your accumulated statistics are not reset. Task domains whose native dependencies fail to build are hidden from the dropdown rather than breaking the app; the startup log lists any that were dropped. Configurable via Space variables: `WAV_EVERY`, `WAV_HORIZON`, `WAV_HORIZON_LONG`, `GPU_DURATION`, `WM_VARIANT`.