Add trained contextual action candidate with browser evaluation and explicit real-web limits
c8e5620 verified |
Download README.md from devildasdf/devils-agent: direct link, hf CLI and curl.
- Browser
- Download file 5.64 kB
-
https://huggingface.co/devildasdf/devils-agent/resolve/main/README.md
- Command line
-
hf download hf://devildasdf/devils-agent/README.md
-
curl -L -o README.md https://huggingface.co/devildasdf/devils-agent/resolve/main/README.md
5.64 kB
| language: | |
| - en | |
| tags: | |
| - browser-agent | |
| - cpu | |
| - experimental | |
| - custom-model | |
| library_name: baim | |
| # Devils Agent / BAIM — experimental research checkpoints | |
| New experimental **v004 contextual checkpoint**: novel-wording fixture match improved from 55.83% to 83.96%; actual browser fixtures from 74/120 to 98/120. Real-web diagnostic remains 1/46. [Results and limitations](docs/CONTEXTUAL_V004.md). | |
| Runtime v0.0.2: **2.38x faster median synthetic CPU prediction**, bounded microbatch inference, and unchanged checkpoint weights. See [measured results and limits](docs/RUNTIME_V002.md). | |
| **Research prototype, not a production-ready general browser agent.** Five custom | |
| checkpoints are stored under `models/`. They require the accompanying Python code; | |
| this repository is not a standard Transformers `AutoModel` or hosted-inference | |
| package. No production champion has been selected. | |
| The checkpoints were trained on generated single-step click/type/select fixtures. | |
| The mean encoder matched only 1 of 46 scorable action/target pairs in a small | |
| Mind2Web diagnostic. Target Linux VPS validation is unfinished. Do not interpret | |
| the high synthetic scores as real-world browser reliability. | |
| No code/model license grant has been selected for this release. Third-party | |
| pretrained weights and raw Mind2Web data are not bundled. See `docs/RESEARCH.md` | |
| and `docs/PRETRAINED.md` for source attribution and external-model provenance. | |
| CPU-first browser action research in progress. **Small policies have been trained, | |
| but there is no validated production checkpoint.** The project includes a browser | |
| runtime, synthetic training pipeline, learned action/pointer policies and measured | |
| evaluations. Real-data transfer is poor; this is not yet a general browser agent. | |
| ## Run | |
| Python 3.12 or newer: | |
| ```sh | |
| python -m venv .venv | |
| # Linux: source .venv/bin/activate | |
| # Windows PowerShell: .venv\Scripts\Activate.ps1 | |
| python -m pip install -e ".[browser]" | |
| python -m playwright install chromium | |
| python -m unittest discover -s tests -v | |
| python -m baim.bench_runtime --output reports/runtime-baseline.json | |
| ``` | |
| The current development host uses `py -3.12` in place of `python` without a virtual | |
| environment for the runtime-only tests. Training uses a virtual environment at | |
| `../../work/baim-venv`. Tests launch isolated headless Chromium contexts and local fixtures. | |
| ## Train and compare | |
| ```sh | |
| python -m pip install torch --index-url https://download.pytorch.org/whl/cpu | |
| python -m pip install -e ".[training,browser]" numpy | |
| python -m baim.synthetic | |
| python -m baim.train --epochs 16 --output models/v000-mean | |
| python -m baim.train --epochs 16 --no-lexical --output models/v001-no-lexical | |
| python -m baim.train --epochs 16 --encoder gru --output models/v002-gru | |
| python -m baim.train --epochs 16 --encoder transformer --output models/v003-transformer | |
| python -m baim.experiment_suite | |
| python -m baim.evaluate_policy --limit 120 | |
| python -m baim.evaluate_policy --data datasets/synthetic-v1/novel_wording.jsonl --output reports/policy-browser-novel-v000.json | |
| python -m baim.evaluate_policy --baseline --data datasets/synthetic-v1/novel_wording.jsonl --output reports/baseline-browser-novel.json | |
| ``` | |
| Mean/GRU/Transformer policies use Hugging Face `PyTorchModelHubMixin` checkpoint | |
| serialization. Nothing is uploaded automatically. FP32 checkpoints use safetensors. | |
| The quantization harness quantizes Linear layers and tests restricted state-dict | |
| reloading; embeddings and recurrent/attention encoders remain FP32. | |
| See [DATASET.md](docs/DATASET.md) for split design and known shortcuts, | |
| [ACCEPTANCE.md](docs/ACCEPTANCE.md) for the full remaining scope, and | |
| `requirements-observed.txt` for package versions measured on this Windows host. | |
| ## Current boundaries | |
| Model output uses a compact action opcode plus a JSON array (for example | |
| `C["e17"]` and `T["e4","hello"]`). It cannot supply selectors or executable JS. | |
| The host supplies the task authority, permission callback and completion verifier. | |
| Task tickets bind session, task epoch, action sequence and observation revision. | |
| References resolve to retained DOM nodes, rechecked before interaction. The adapter | |
| includes frames and open shadow roots. Browser actionability checks still apply. | |
| Recorded telemetry contains metadata and keyed hashes; it excludes goal text, | |
| typed values, raw observations and extracted content. This is **not yet sufficient | |
| training data**. The key is a separate local file; protect both files and set an | |
| appropriate retention policy. Sensitive-target detection is incomplete and must | |
| not be mistaken for comprehensive PII detection. | |
| This is not a browser security sandbox. The permission callback must enforce the | |
| deployment's trusted action policy. Network isolation, redirect restrictions, | |
| download policy and rich redacted trajectories remain to be implemented. A page | |
| can change between validation and interaction; hostile timing attacks are not | |
| solved by node handles. Accessible names are an approximation, not full ARIA | |
| accessible-name computation. Closed shadow roots and canvas require fallback. | |
| The learned baseline covers CLICK, TYPE and SELECT, with literal copying from | |
| one quoted user-goal value. It has no general planner, history model or reliable | |
| unsupported-task detector. Validation temperature scaling fails under distribution | |
| shift; confidence is not a security boundary. Keep it on isolated research fixtures. | |
| FINISH requires a host verifier bound to the current task. The local tests supply | |
| fixture-specific verifiers; arbitrary user-goal completion is still unresolved. | |
| See [docs/STATUS.md](docs/STATUS.md) and the full [requirements](docs/requirements.txt). | |