devils-agent / docs /CONTEXTUAL_V004.md
devildasdf's picture
Add trained contextual action candidate with browser evaluation and explicit real-web limits
c8e5620 verified
|
Raw History Blame Contribute Delete
2.13 kB

v004: target-aware action classification

The previous action head read only the goal. The new optional contextual_action head also reads a soft attention-weighted representation of the DOM candidates, using the learned pointer logits. This is trained end-to-end, without runtime phrase-to-action rules. Existing checkpoints load with the feature disabled.

Training: unchanged synthetic-v1 train split, 2,400 examples; 16 epochs; seed 1729; mean encoder, width 64, 132,996 parameters. Selection uses validation joint accuracy only (first best epoch wins ties). Temperatures fit validation only. The existing novel-wording evaluation is not training data, but was previously inspected during development, so it is not a fresh blind benchmark.

Results:

Check Previous mean v004
Novel-wording action + target match, 480 fixtures 55.83% 83.96%
Actual Chromium novel fixtures, 120 tasks 74/120 98/120
Standard synthetic test match 100% 100%
Small Mind2Web diagnostic 1/46 1/46

CPU encoding plus neural forward pass: 1.103 ms median, 1.647 ms p95 on Windows with two Torch threads. Browser policy prediction measured 2.046 ms median with Chromium running. These different timing scopes must not be compared as equivalent. Checkpoint size is 532,696 bytes. Reports include platform and measurement scope.

Usage:

from baim.policy import LearnedPolicy
policy = LearnedPolicy('models/v004-contextual')
decision = policy.predict(goal, observed_state, authority_ticket)

Reproduce training:

python -m baim.train --contextual-action --epochs 16 --output models/v004-contextual

This candidate is not production-promoted. Real-site grounding, calibration under distribution shift, long goals, and general planning remain unresolved. It still predicts only click/type/select and copies one quoted literal for values. The live AegisVision default remains unchanged. Independent review and approval must remain in place. Reports: cpu-v004.json, browser-v004-novel.json, and mind2web-v004.json under reports/, plus the checkpoint training report.