Von Browser-Act
Von 1.2, fine-tuned into a single-pass browser action policy. Given the goal, the page and a finite list of candidate actions, it picks one β a control and what to do with it, or a move like scroll, wait, done or blocked β in one encoder pass, with a calibrated probability. It is the local "System 1" behind the Act tool of Lychee AI, a Chrome side-panel agent, where it runs inside the browser on WebGPU (the 8-bit ONNX export at this repo's root, 616 MB) and decides a step in about 400 ms while the chat model waits.
It is a fine-tune of wfzyx/von (Von 1.2: a 395M ModernBERT-large encoder with an Option-Marker head, Apache-2.0). Nothing about the architecture changed; what changed is the weights and the question they are trained to answer.
What it does
The browser harness, not the model, owns the action space. It reads the page, builds the finite set of valid (control, action) pairs plus the global moves that are possible right now, and asks the model one question:
Goal: add the blue Aurora mug to the cart
Page: Aurora Home β Shop (http://127.0.0.1:52439)
State: scroll down: yes; scroll up: no
Provided values: none
Recent actions: none
with options like
[8] button "Add Aurora mug (blue) to cart" β click
[9] button "Add Aurora mug (green) to cart" β click
[2] link "Cart" β click
[14] combobox "Sort by" (currently "Featured") β click
SCROLL_DOWN Scroll down: what is needed next is below the visible area.
WAIT The page is still loading or changing; wait for it to settle.
DONE The goal is already achieved on this page.
BLOCKED The goal cannot be progressed here.
The answer is one option id with a calibrated probability over all of them. For a field, a second small question picks which of the caller-supplied values goes in. The model never produces text, a selector, a URL or a navigation; the harness validates the answer against the set it offered before anything runs, and every committing step still goes through a human approval card. Page text is untrusted input to a classifier, and the closed set is what bounds an injection to a wrong non-committing click.
This design follows browser-use/jev-ultrafast, which scopes TypeSafe's Jev the same way; this model reproduces its "one request per decision" with open weights that run client-side.
Numbers
Held-out validation (363 rows from 12 sites never seen in training):
| head | rows | top-1 | coverage at the 0.6 gate | accuracy on the covered |
|---|---|---|---|---|
| all | 363 | 0.939 | 0.939 | 0.971 |
| element options (a control + an action) | 197 | 0.919 | 0.914 | 0.972 |
| DONE | 118 | 1.000 | 1.000 | 1.000 |
| SCROLL_DOWN | 22 | 0.818 | 0.864 | 0.789 |
| value (which supplied value goes in the field) | 21 | 1.000 | 1.000 | 1.000 |
Base Von 1.2 scores 0.240 on the same rows: the question format is new to it.
Real browser (the Lychee eval harness, Chrome for Testing 151, WebGPU on an M4 Max, 54 goals on three fixture pages, 2026-09-30):
| this model | the chat model (gemma-4-26b-a4b, same candidate lists) | |
|---|---|---|
| grounding top-1 | 50/54 = 0.926 | 51/54 = 0.944 |
| passes per click step | 1 | 1 completion |
| p50 per decision | 400 ms (817 ms for 47 options in one pass) | 1.7 s |
| WebGPU vs fp32 parity (q8) | max |Ξp| 0.0098, 0 argmax flips |
End to end, on a cookie-wall-then-add-to-cart task it performs both clicks in one call and stops on its own; a forced-arm run finished in 29 s against 38 s for the chat model doing every step itself.
General decision benchmarks (the fine-tune did not damage Von's general ability; CPU, the reference von-sdk scorer, same seeds):
| benchmark | base Von 1.2 | this model | Jev 1.13 (closed) |
|---|---|---|---|
| jabr v2, 866 items | 0.724 | 0.718 | 0.964 |
| jabr v2 ECE | 0.145 | 0.060 | |
| JevBench public hard, 111 items | 0.368 | 0.514 | |
| JevBench public "original", choice items, 36 | 0.722 | 0.500 | |
| DecisionBench, 492-row stratified subsample | 0.490 | 0.520 | 0.720 (full suite) |
Per-item latency with the SDK: p50 0.126 s on CPU, 0.026 s on Apple MPS for jabr/JevBench-sized items; MPS matches CPU with zero argmax flips.
Files
| path | what | for |
|---|---|---|
model_q8.onnx, tokenizer.json, tokenizer_config.json, von.json, manifest.json |
the browser export: weight-only 8-bit (block-16 asymmetric on the attention-output and MLP matmuls; Wqkv, embeddings and the scorer stored fp16), fp32 activations, so no shader-f16 is needed; ~0.9 GB of device memory |
onnxruntime-web on WebGPU. manifest.json lists every file with its SHA-256 for a hash-verified download; von.json carries the special-token ids, the 4096-token client cap and the calibration temperature |
von-golden.json |
fp32 reference logits for a few packed inputs | a parity test for any port |
checkpoint/ |
the PyTorch checkpoint in Von's own layout: option_marker.pt (the fine-tuned state dict), model.safetensors + config.json (the base backbone files the SDK loads first), tokenizer, marker_calibration.json, train_meta.json |
von-sdk, further fine-tuning, von serve |
Use it
In the browser (onnxruntime-web). Pack "{instructions} {state}" [SEP] [MASK] option0 [MASK] option1 β¦, tokenize, and run with four int64 inputs: input_ids, option_ids (β1 for prefix tokens, the option index for each option's tokens), position_ids (each option's positions restart at the prefix length), marker_positions (the index of each [MASK]). The output logits has one entry per marker; divide by the temperature in von.json and softmax. Options attend only to the prefix and themselves, so the answer is order-invariant and you may pack as many questions' options after one shared prefix as fit. The reference implementation is src/system1/{von,worker,vonPolicy}.ts in the Lychee repository, with the exact renderers that produced the training rows.
With the Von SDK (Python).
pip install "von-sdk @ git+https://github.com/wfzyx/von"
hf download ks-ang/von-browser-act --include "checkpoint/*" --local-dir ./von-browser-act
cd ./von-browser-act/checkpoint && von serve --device cpu --port 8010 # TypeSafe-compatible /v1/systemone
On Apple GPUs serve with --device mps but serialize requests: concurrent requests crash the MPS backend (a Metal "command encoder already encoding" assertion, measured with DecisionBench's 16-way client).
In Lychee AI. Pin this repo's commit and the manifest hash in src/system1/weightsManifest.ts, build, and turn on Settings β General β Experimental β "Fast page control"; the extension downloads and hash-verifies the export into Cache Storage and offers Act inside a granted page-control session.
Training
- Data. 4,921 episodes over ~190 public sites, collected read-only (snapshots at three scroll offsets with the extension's own DOM walker) and labelled by two local teacher models (
google/gemma-4-26b-a4b,qwen3.8-27b) in hindsight: a teacher writes goals each fulfilled by one visible control, the other teacher must ground the goal back to the same control blind, and disagreements are dropped. Moves are labelled structurally (a goal for a control further down β SCROLL_DOWN; a step already in the recent actions with the field showing its value β DONE; a missing value β BLOCKED). Rendered through the shipped harness code into 2,743 training rows (single-round element goals, DONE, SCROLL_DOWN, BLOCKED, tournament chunk/final rows above 120 options, value rows); the split is by site. - Recipe. Full fine-tune of Von 1.2 (
wfzyx/von@5df8185), listwise cross-entropy + 0.5Β·Brier over the option markers with independent-option masking, lr 1e-5 (head 5e-5), 3 epochs = 396 optimizer steps, gradient checkpointing, 72.7 min on an M4 Max; best checkpoint at step 250. Calibration is a single refit temperature (1.389); the base model's input-conditioned calibration map is not used. - Export.
scripts/export-von/export.pyin the Lychee repository: ONNX with the order-invariant attention mask built inside the graph, 8-bit weight-only quantization gated on parity with the fp32 reference (max |Ξp| β€ 0.03; this export: 0.0098).
Limitations
- Dropdowns and long forms. Only nine training rows are
SELECT_OPTION; on a real form the model over-selects a dropdown and hands the value question a key that matches no option. The harness now excludes a refused pair and moves on, but the data gap is real. Multi-field trajectories are also thin, so on long forms it tends to stop with a low-confidenceDONEorBLOCKEDafter a few fields. - Confidence scales with the option count. The calibrated top over 47 options sits at 0.36β0.50 for some correct answers; the harness uses a two-band gate (act freely at 0.6, tentatively at 0.4 when the step is not committing).
- No page text in the state. The model sees the goal, the page title, the controls and its own recent actions, not the page's prose, so a "done" that only shows as a toast or a results list is invisible to it; the chat model verifies after
done. - English, desktop pages, the extension's control vocabulary. It expects the exact line format above (
[id] role "name" (state) β action); other renderings are out of distribution. - Evaluated on three fixture pages and a handful of public benchmarks, not a broad web benchmark. A valid choice can still be wrong; keep a verifier and a gate around it.
License and attribution
Apache-2.0, like the base model. Von is by wfzyx (wfzyx/von, Apache-2.0); this repository redistributes its backbone files unchanged under checkpoint/ alongside the fine-tuned head state. The training, export and evaluation code lives in the Lychee AI repository (scripts/von-finetune, scripts/export-von, scripts/system1-eval).
@misc{vonbrowseract2026,
title = {Von Browser-Act: Von 1.2 fine-tuned as a single-pass browser action policy},
author = {Ang, Kah Shin},
year = {2026},
url = {https://huggingface.co/ks-ang/von-browser-act}
}