laya-web-agent

A fast decision model for browser agents. Given a user's goal and the current page, it picks the next step in a single forward pass, without generating any text:

  • operation: CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN, WAIT, DONE or BLOCKED
  • target: which numbered control on the page to act on

The input is the page state a browser agent sends: title, visible text, a numbered list of the controls on screen with their current values, and the recent actions.

Architecture Laya decision model on a ModernBERT-large encoder (28 layers, 1024 hidden)
Parameters 421M
Context 2048 tokens (up to 512 for the question and its options)
Output typed answers with probabilities, one forward pass

Usage

import laya
from huggingface_hub import snapshot_download

agent = laya.load(snapshot_download("abedinia/laya-web-agent"))

The input format is jev_ultrafast's: the state plus typed operation and target questions. rl_agent_config.json holds the lengths the model was trained with (max_len 2048, head_max_len 512). Keep them. With shorter lengths, field values and history get cut off.

Results

Real websites: Mind2Web official test splits

For each step, the top 25 candidate elements come from the MindAct DeBERTa ranker (its scores ship with the dataset), listed in page order. A step counts as a miss if the ranker drops the correct element, or if the step couldn't be converted (about 1%).

split steps element operation step success step success, macro per task
cross-task (new tasks, known sites) 2,094 35.0% 85.8% 26.1% 29.1%
cross-website (new sites) 1,373 27.1% 82.1% 18.6% 21.4%
cross-domain (new domains) 5,911 27.8% 84.5% 19.4% 21.7%
  • element: the right control.
  • operation: the right action (CLICK, TYPE or SELECT). Typed values aren't scored, because the agent fills them in separately.
  • step success: both right.
  • The shortlist caps element accuracy: the ranker keeps the correct element in its top 25 on 82%, 77% and 78% of steps.

Structured pages: synthetic test

500 held-out episodes (2,381 decisions) of 17 page types, generated from seeds training never used and played in real Chrome:

per step, both questions 98.5%
operation 97.8%
target 99.4%
episodes with every step right 87.7% (436 of 497)

It's at 100% on checkboxes, wizards, contact messages, already-completed pages and impossible tasks (BLOCKED, 12/12). The weakest page types are carts (86.5% on operation) and recovering after a mistyped form value (90.2%). Most remaining mistakes are CLICK-or-SCROLL_DOWN calls when a control sits at the very bottom of the screen.

Training

Data: 33,438 decisions.

  • 26,234 from structured synthetic pages, 5,648 episodes. Each was played in real headless Chrome, and the correct answer at every step comes from checking the live page, not from a model. The 17 page types are forms, signup, address, contact message, search, autocomplete, filters, cart, date picker, two-step wizard, dropdown, checkbox, radio, navigation, mixed bookings, already-completed pages and impossible tasks. Situations covered:
    • recovering after a mistake (early submit, wrong value, wrong field, wrong option),
    • scrolling to controls below the fold,
    • waiting through loading states,
    • ignoring instructions injected into the page text,
    • randomised layout and wording.
  • 7,204 steps from real websites: the Mind2Web training split, converted to the same input format.

Labels: CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN (1,552), WAIT (421), DONE (5,474) and BLOCKED (143).

Setup:

  • 3 epochs over 50,164 training items.
  • The top 16 of 28 encoder layers plus the decision head were trained (222.6M parameters).
  • AdamW, lr 3e-5, effective batch 16, bf16 with gradient checkpointing, 2048-token context.
  • About 11 hours on one RTX 4070 laptop GPU (8 GB).
  • The holdout (15% of episodes, split by episode) peaked at 0.890 in the final epoch.

Limitations

  • Real websites: about 1 step in 5 is fully right. It usually picks the right kind of action, but on crowded pages it often picks the wrong one of many similar controls. That isn't enough for long real-world tasks on its own. It works best on structured pages, or as a fast first pass that hands off to a larger model when its confidence is low.
  • Assumed formats: the representation of scroll and wait steps in the action history, and of radio buttons and textareas, follows the agent's format as closely as we could reconstruct it.
  • English: all training data is English.
  • DONE without seeing the confirmation: on long pages the confirmation can be below the fold, and the model answers DONE right after the final submit.

Credits

  • Fine-tuned from Quantum08/laya-browser-mind2web (Apache-2.0), which builds on convaiinnovations/laya.
  • Mind2Web: Deng et al., Mind2Web: Towards a Generalist Agent for the Web, NeurIPS 2023 Datasets and Benchmarks Track, CC BY 4.0. The Mind2Web test set was used only for evaluation and isn't redistributed here.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abedinia/laya-web-agent

Finetuned
(1)
this model

Dataset used to train abedinia/laya-web-agent