LGWM: Latent GUI World Model

Weights for Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents.

Jiaming Zhang, Xuan Wang, Fuyao Zhang, Yang Cao, Lingjuan Lyu, Wei Yang Bryan Lim

Project page · Code · Data · Paper coming soon

LGWM compares predicted and observed screens in one representation space. The same login screen is legitimate after one action and a hijack after another.

LGWM is an action-conditioned world model that predicts the next screen in representation space and checks it against the screen that actually appears. One cosine distance verifies every step of a GUI agent, with no judge model and no attack labels.

Key results

  • 0.987 AUC on RSWT-Bench with a training-free integrity score, within one AUC point of Gemini 3.7 Flash (0.995) and GPT-5.6 Luna (0.988).
  • 17 ms per decision on one A100. World models that render the future (gWorld, Code2World) score about ten AUC points lower and need over 90 s per transition.
  • 0.954 AUC at separating harmful from benign violations with a linear head on the prediction residual, where prompted VLMs score between 0.51 and 0.61.
  • 261M parameters in training, 173M at inference.

Files

Directory Model
lgwm-online/ World model: ViT-B/14 screen encoder, action encoder and 12-layer predictor (173M parameters)
lgwm-harm-head/ Linear harm head on the prediction residual

Usage

git clone https://github.com/jiamingzhang94/lgwm && cd lgwm
pip install -c requirements-tested.txt -e '.[text]'
hf download jiamingzz/lgwm --local-dir weights

echo '{"action_type": "tap", "x": 0.5, "y": 0.5}' > action.json
python tools/score_transition.py --before before.png --after after.png \
  --action action.json --harm-head weights/lgwm-harm-head

The command prints integrity, which is 1 − cos(predicted, observed), and harm_score. Higher values indicate a less expected transition and a higher predicted hijack probability.

Training

Self-supervised on 1.85M real GUI transitions from Android in the Wild, AndroidControl, GUIOdyssey, AMEX and MiniWoB++, starting from DINOv2 ViT-B/14: 120,000 updates at batch size 512 on four A100 GPUs. The harm head is fit on the 6,786-example supervised split of RSWT-Bench.

Citation

@misc{zhang2026lgwm,
  title  = {Right Screen, Wrong Transition:
            World Models as Verifiers for GUI Agents},
  author = {Zhang, Jiaming and Wang, Xuan and Zhang, Fuyao and
            Cao, Yang and Lyu, Lingjuan and Lim, Wei Yang Bryan},
  year   = {2026},
  howpublished = {\url{https://github.com/jiamingzhang94/lgwm}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jiamingzz/lgwm

Finetuned
(108)
this model

Dataset used to train jiamingzz/lgwm

Evaluation results