Download docs/CONTEXTUAL_V004.md from devildasdf/devils-agent: direct link, hf CLI and curl.
- Browser
- Download file 2.13 kB
-
https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/CONTEXTUAL_V004.md
- Command line
-
hf download hf://devildasdf/devils-agent/docs/CONTEXTUAL_V004.md
-
curl -L -o CONTEXTUAL_V004.md https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/CONTEXTUAL_V004.md
v004: target-aware action classification
The previous action head read only the goal. The new optional contextual_action
head also reads a soft attention-weighted representation of the DOM candidates,
using the learned pointer logits. This is trained end-to-end, without runtime
phrase-to-action rules. Existing checkpoints load with the feature disabled.
Training: unchanged synthetic-v1 train split, 2,400 examples; 16 epochs; seed 1729; mean encoder, width 64, 132,996 parameters. Selection uses validation joint accuracy only (first best epoch wins ties). Temperatures fit validation only. The existing novel-wording evaluation is not training data, but was previously inspected during development, so it is not a fresh blind benchmark.
Results:
| Check | Previous mean | v004 |
|---|---|---|
| Novel-wording action + target match, 480 fixtures | 55.83% | 83.96% |
| Actual Chromium novel fixtures, 120 tasks | 74/120 | 98/120 |
| Standard synthetic test match | 100% | 100% |
| Small Mind2Web diagnostic | 1/46 | 1/46 |
CPU encoding plus neural forward pass: 1.103 ms median, 1.647 ms p95 on Windows with two Torch threads. Browser policy prediction measured 2.046 ms median with Chromium running. These different timing scopes must not be compared as equivalent. Checkpoint size is 532,696 bytes. Reports include platform and measurement scope.
Usage:
from baim.policy import LearnedPolicy
policy = LearnedPolicy('models/v004-contextual')
decision = policy.predict(goal, observed_state, authority_ticket)
Reproduce training:
python -m baim.train --contextual-action --epochs 16 --output models/v004-contextual
This candidate is not production-promoted. Real-site grounding, calibration
under distribution shift, long goals, and general planning remain unresolved.
It still predicts only click/type/select and copies one quoted literal for values.
The live AegisVision default remains unchanged. Independent review and approval
must remain in place. Reports: cpu-v004.json, browser-v004-novel.json, and
mind2web-v004.json under reports/, plus the checkpoint training report.