Add trained contextual action candidate with browser evaluation and explicit real-web limits
c8e5620 verified |
Download docs/CONTEXTUAL_V004.md from devildasdf/devils-agent: direct link, hf CLI and curl.
- Browser
- Download file 2.13 kB
-
https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/CONTEXTUAL_V004.md
- Command line
-
hf download hf://devildasdf/devils-agent/docs/CONTEXTUAL_V004.md
-
curl -L -o CONTEXTUAL_V004.md https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/CONTEXTUAL_V004.md
2.13 kB
| # v004: target-aware action classification | |
| The previous action head read only the goal. The new optional `contextual_action` | |
| head also reads a soft attention-weighted representation of the DOM candidates, | |
| using the learned pointer logits. This is trained end-to-end, without runtime | |
| phrase-to-action rules. Existing checkpoints load with the feature disabled. | |
| Training: unchanged synthetic-v1 train split, 2,400 examples; 16 epochs; seed 1729; | |
| mean encoder, width 64, 132,996 parameters. Selection uses validation joint accuracy | |
| only (first best epoch wins ties). Temperatures fit validation only. The existing | |
| novel-wording evaluation is not training data, but was previously inspected during | |
| development, so it is not a fresh blind benchmark. | |
| Results: | |
| | Check | Previous mean | v004 | | |
| |---|---:|---:| | |
| | Novel-wording action + target match, 480 fixtures | 55.83% | 83.96% | | |
| | Actual Chromium novel fixtures, 120 tasks | 74/120 | 98/120 | | |
| | Standard synthetic test match | 100% | 100% | | |
| | Small Mind2Web diagnostic | 1/46 | 1/46 | | |
| CPU encoding plus neural forward pass: 1.103 ms median, 1.647 ms p95 on Windows | |
| with two Torch threads. Browser policy prediction measured 2.046 ms median with | |
| Chromium running. These different timing scopes must not be compared as equivalent. | |
| Checkpoint size is 532,696 bytes. Reports include platform and measurement scope. | |
| Usage: | |
| ```python | |
| from baim.policy import LearnedPolicy | |
| policy = LearnedPolicy('models/v004-contextual') | |
| decision = policy.predict(goal, observed_state, authority_ticket) | |
| ``` | |
| Reproduce training: | |
| ```sh | |
| python -m baim.train --contextual-action --epochs 16 --output models/v004-contextual | |
| ``` | |
| This candidate is **not production-promoted**. Real-site grounding, calibration | |
| under distribution shift, long goals, and general planning remain unresolved. | |
| It still predicts only click/type/select and copies one quoted literal for values. | |
| The live AegisVision default remains unchanged. Independent review and approval | |
| must remain in place. Reports: `cpu-v004.json`, `browser-v004-novel.json`, and | |
| `mind2web-v004.json` under `reports/`, plus the checkpoint training report. | |