devils-agent / docs /CONTEXTUAL_V004.md
devildasdf's picture
Add trained contextual action candidate with browser evaluation and explicit real-web limits
c8e5620 verified
|
Raw History Blame Contribute Delete
2.13 kB
# v004: target-aware action classification
The previous action head read only the goal. The new optional `contextual_action`
head also reads a soft attention-weighted representation of the DOM candidates,
using the learned pointer logits. This is trained end-to-end, without runtime
phrase-to-action rules. Existing checkpoints load with the feature disabled.
Training: unchanged synthetic-v1 train split, 2,400 examples; 16 epochs; seed 1729;
mean encoder, width 64, 132,996 parameters. Selection uses validation joint accuracy
only (first best epoch wins ties). Temperatures fit validation only. The existing
novel-wording evaluation is not training data, but was previously inspected during
development, so it is not a fresh blind benchmark.
Results:
| Check | Previous mean | v004 |
|---|---:|---:|
| Novel-wording action + target match, 480 fixtures | 55.83% | 83.96% |
| Actual Chromium novel fixtures, 120 tasks | 74/120 | 98/120 |
| Standard synthetic test match | 100% | 100% |
| Small Mind2Web diagnostic | 1/46 | 1/46 |
CPU encoding plus neural forward pass: 1.103 ms median, 1.647 ms p95 on Windows
with two Torch threads. Browser policy prediction measured 2.046 ms median with
Chromium running. These different timing scopes must not be compared as equivalent.
Checkpoint size is 532,696 bytes. Reports include platform and measurement scope.
Usage:
```python
from baim.policy import LearnedPolicy
policy = LearnedPolicy('models/v004-contextual')
decision = policy.predict(goal, observed_state, authority_ticket)
```
Reproduce training:
```sh
python -m baim.train --contextual-action --epochs 16 --output models/v004-contextual
```
This candidate is **not production-promoted**. Real-site grounding, calibration
under distribution shift, long goals, and general planning remain unresolved.
It still predicts only click/type/select and copies one quoted literal for values.
The live AegisVision default remains unchanged. Independent review and approval
must remain in place. Reports: `cpu-v004.json`, `browser-v004-novel.json`, and
`mind2web-v004.json` under `reports/`, plus the checkpoint training report.