Von Browser-Act

Von 1.2, fine-tuned into a single-pass browser action policy. Given the goal, the page and a finite list of candidate actions, it picks one β€” a control and what to do with it, or a move like scroll, wait, done or blocked β€” in one encoder pass, with a calibrated probability. It is the local "System 1" behind the Act tool of Lychee AI, a Chrome side-panel agent, where it runs inside the browser on WebGPU (the 8-bit ONNX export at this repo's root, 616 MB) and decides a step in about 400 ms while the chat model waits.

It is a fine-tune of wfzyx/von (Von 1.2: a 395M ModernBERT-large encoder with an Option-Marker head, Apache-2.0). Nothing about the architecture changed; what changed is the weights and the question they are trained to answer.

What it does

The browser harness, not the model, owns the action space. It reads the page, builds the finite set of valid (control, action) pairs plus the global moves that are possible right now, and asks the model one question:

Goal: add the blue Aurora mug to the cart
Page: Aurora Home β€” Shop (http://127.0.0.1:52439)
State: scroll down: yes; scroll up: no
Provided values: none
Recent actions: none

with options like

[8]  button "Add Aurora mug (blue) to cart" β†’ click
[9]  button "Add Aurora mug (green) to cart" β†’ click
[2]  link "Cart" β†’ click
[14] combobox "Sort by" (currently "Featured") β†’ click
SCROLL_DOWN  Scroll down: what is needed next is below the visible area.
WAIT         The page is still loading or changing; wait for it to settle.
DONE         The goal is already achieved on this page.
BLOCKED      The goal cannot be progressed here.

The answer is one option id with a calibrated probability over all of them. For a field, a second small question picks which of the caller-supplied values goes in. The model never produces text, a selector, a URL or a navigation; the harness validates the answer against the set it offered before anything runs, and every committing step still goes through a human approval card. Page text is untrusted input to a classifier, and the closed set is what bounds an injection to a wrong non-committing click.

This design follows browser-use/jev-ultrafast, which scopes TypeSafe's Jev the same way; this model reproduces its "one request per decision" with open weights that run client-side.

Numbers

Held-out validation (363 rows from 12 sites never seen in training):

head rows top-1 coverage at the 0.6 gate accuracy on the covered
all 363 0.939 0.939 0.971
element options (a control + an action) 197 0.919 0.914 0.972
DONE 118 1.000 1.000 1.000
SCROLL_DOWN 22 0.818 0.864 0.789
value (which supplied value goes in the field) 21 1.000 1.000 1.000

Base Von 1.2 scores 0.240 on the same rows: the question format is new to it.

Real browser (the Lychee eval harness, Chrome for Testing 151, WebGPU on an M4 Max, 54 goals on three fixture pages, 2026-09-30):

this model the chat model (gemma-4-26b-a4b, same candidate lists)
grounding top-1 50/54 = 0.926 51/54 = 0.944
passes per click step 1 1 completion
p50 per decision 400 ms (817 ms for 47 options in one pass) 1.7 s
WebGPU vs fp32 parity (q8) max |Ξ”p| 0.0098, 0 argmax flips

End to end, on a cookie-wall-then-add-to-cart task it performs both clicks in one call and stops on its own; a forced-arm run finished in 29 s against 38 s for the chat model doing every step itself.

General decision benchmarks (the fine-tune did not damage Von's general ability; CPU, the reference von-sdk scorer, same seeds):

benchmark base Von 1.2 this model Jev 1.13 (closed)
jabr v2, 866 items 0.724 0.718 0.964
jabr v2 ECE 0.145 0.060
JevBench public hard, 111 items 0.368 0.514
JevBench public "original", choice items, 36 0.722 0.500
DecisionBench, 492-row stratified subsample 0.490 0.520 0.720 (full suite)

Per-item latency with the SDK: p50 0.126 s on CPU, 0.026 s on Apple MPS for jabr/JevBench-sized items; MPS matches CPU with zero argmax flips.

Files

path what for
model_q8.onnx, tokenizer.json, tokenizer_config.json, von.json, manifest.json the browser export: weight-only 8-bit (block-16 asymmetric on the attention-output and MLP matmuls; Wqkv, embeddings and the scorer stored fp16), fp32 activations, so no shader-f16 is needed; ~0.9 GB of device memory onnxruntime-web on WebGPU. manifest.json lists every file with its SHA-256 for a hash-verified download; von.json carries the special-token ids, the 4096-token client cap and the calibration temperature
von-golden.json fp32 reference logits for a few packed inputs a parity test for any port
checkpoint/ the PyTorch checkpoint in Von's own layout: option_marker.pt (the fine-tuned state dict), model.safetensors + config.json (the base backbone files the SDK loads first), tokenizer, marker_calibration.json, train_meta.json von-sdk, further fine-tuning, von serve

Use it

In the browser (onnxruntime-web). Pack "{instructions} {state}" [SEP] [MASK] option0 [MASK] option1 …, tokenize, and run with four int64 inputs: input_ids, option_ids (βˆ’1 for prefix tokens, the option index for each option's tokens), position_ids (each option's positions restart at the prefix length), marker_positions (the index of each [MASK]). The output logits has one entry per marker; divide by the temperature in von.json and softmax. Options attend only to the prefix and themselves, so the answer is order-invariant and you may pack as many questions' options after one shared prefix as fit. The reference implementation is src/system1/{von,worker,vonPolicy}.ts in the Lychee repository, with the exact renderers that produced the training rows.

With the Von SDK (Python).

pip install "von-sdk @ git+https://github.com/wfzyx/von"
hf download ks-ang/von-browser-act --include "checkpoint/*" --local-dir ./von-browser-act
cd ./von-browser-act/checkpoint && von serve --device cpu --port 8010   # TypeSafe-compatible /v1/systemone

On Apple GPUs serve with --device mps but serialize requests: concurrent requests crash the MPS backend (a Metal "command encoder already encoding" assertion, measured with DecisionBench's 16-way client).

In Lychee AI. Pin this repo's commit and the manifest hash in src/system1/weightsManifest.ts, build, and turn on Settings β†’ General β†’ Experimental β†’ "Fast page control"; the extension downloads and hash-verifies the export into Cache Storage and offers Act inside a granted page-control session.

Training

  • Data. 4,921 episodes over ~190 public sites, collected read-only (snapshots at three scroll offsets with the extension's own DOM walker) and labelled by two local teacher models (google/gemma-4-26b-a4b, qwen3.8-27b) in hindsight: a teacher writes goals each fulfilled by one visible control, the other teacher must ground the goal back to the same control blind, and disagreements are dropped. Moves are labelled structurally (a goal for a control further down β†’ SCROLL_DOWN; a step already in the recent actions with the field showing its value β†’ DONE; a missing value β†’ BLOCKED). Rendered through the shipped harness code into 2,743 training rows (single-round element goals, DONE, SCROLL_DOWN, BLOCKED, tournament chunk/final rows above 120 options, value rows); the split is by site.
  • Recipe. Full fine-tune of Von 1.2 (wfzyx/von@5df8185), listwise cross-entropy + 0.5Β·Brier over the option markers with independent-option masking, lr 1e-5 (head 5e-5), 3 epochs = 396 optimizer steps, gradient checkpointing, 72.7 min on an M4 Max; best checkpoint at step 250. Calibration is a single refit temperature (1.389); the base model's input-conditioned calibration map is not used.
  • Export. scripts/export-von/export.py in the Lychee repository: ONNX with the order-invariant attention mask built inside the graph, 8-bit weight-only quantization gated on parity with the fp32 reference (max |Ξ”p| ≀ 0.03; this export: 0.0098).

Limitations

  • Dropdowns and long forms. Only nine training rows are SELECT_OPTION; on a real form the model over-selects a dropdown and hands the value question a key that matches no option. The harness now excludes a refused pair and moves on, but the data gap is real. Multi-field trajectories are also thin, so on long forms it tends to stop with a low-confidence DONE or BLOCKED after a few fields.
  • Confidence scales with the option count. The calibrated top over 47 options sits at 0.36–0.50 for some correct answers; the harness uses a two-band gate (act freely at 0.6, tentatively at 0.4 when the step is not committing).
  • No page text in the state. The model sees the goal, the page title, the controls and its own recent actions, not the page's prose, so a "done" that only shows as a toast or a results list is invisible to it; the chat model verifies after done.
  • English, desktop pages, the extension's control vocabulary. It expects the exact line format above ([id] role "name" (state) β†’ action); other renderings are out of distribution.
  • Evaluated on three fixture pages and a handful of public benchmarks, not a broad web benchmark. A valid choice can still be wrong; keep a verifier and a gate around it.

License and attribution

Apache-2.0, like the base model. Von is by wfzyx (wfzyx/von, Apache-2.0); this repository redistributes its backbone files unchanged under checkpoint/ alongside the fine-tuned head state. The training, export and evaluation code lives in the Lychee AI repository (scripts/von-finetune, scripts/export-von, scripts/system1-eval).

@misc{vonbrowseract2026,
  title  = {Von Browser-Act: Von 1.2 fine-tuned as a single-pass browser action policy},
  author = {Ang, Kah Shin},
  year   = {2026},
  url    = {https://huggingface.co/ks-ang/von-browser-act}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ks-ang/von-browser-act

Finetuned
wfzyx/von
Quantized
(4)
this model