Laya 路 ONNX (fp16) for the browser

An unofficial ONNX export of convaiinnovations/laya, the open-source "System One" decision model by Convai Innovations, made to run client-side with ONNX Runtime Web on WebGPU or WebAssembly. Laya takes a state and typed questions and scores every option in one forward pass. It returns calibrated probabilities, not generated text.

This is the default model of @wexare/laya-web (source), which loads it automatically:

import { LayaWorkerClient } from "@wexare/laya-web";

Files

File What it is
laya_fp16.onnx The model: one self-contained file, 900,153,849 bytes, opset 18
tokenizer/tokenizer.json, tokenizer/tokenizer_config.json The checkpoint's ModernBERT tokenizer, mirrored unchanged
rl_agent_config.json Calibration temperatures and length limits, mirrored unchanged
export.json Export metadata, including the model's sha256
laya_q8.onnx Experimental 8-bit build, WebAssembly only: 632,942,133 bytes, see below
laya_q8.json Its metadata and measured agreement

sha256 of laya_fp16.onnx: b2125c7a1044e78807ee718ff4e65c9892b6b55aeeeee64f31fffffd76871b0a

Graph

inputs   input_ids [B,L] int64, attention_mask [B,L] int64,
         marker_pos [B,K] int64, marker_mask [B,K] bool, qtype [B] int64
outputs  logits [B,K] float32, act_logits [B,2] float32

Batch, sequence and option count are all dynamic. Masked option slots return -1e4. Turning logits into answers (per-option-count temperature, softmax, confidence) follows the reference laya package and is done by the library, not the graph.

Precision

Mixed precision, set up in PyTorch before export:

  • The ModernBERT-large encoder (about 95% of the weights) is in fp16, with every LayerNorm computing in fp32.
  • The decision head (2 transformer layers, option scorer, act head) stays in fp32.
  • Inputs and outputs are the types listed above; no fp16 tensors cross the graph boundary.

Verification

Against the reference laya 0.3.3 package, on 9 fixture cases covering single and multi-question batches, 2 to 14 options, a 600-token truncated state, structured and non-ASCII inputs:

  • fp32 export vs PyTorch: identical decisions, maximum probability difference 0.00000.
  • This fp16 file vs PyTorch fp32 (ONNX Runtime 1.30, CPU): identical decisions, maximum probability difference 0.00068.
  • Through @wexare/laya-web on onnxruntime-web 1.30 (WASM): 9/9 end-to-end parity tests pass: same chosen option and same yes/no verdict, with probabilities within 0.02 of the Python model.

Experimental: 8-bit (laya_q8.onnx)

A smaller build for testing on phones, where the fp16 model can exhaust the browser's memory while loading. Its MatMul weights are quantized to 8 bits with ONNX Runtime's MatMulNBits (block 32, symmetric); everything else stays fp32. WebAssembly only: onnxruntime-web's WebGPU MatMulNBits kernel accepts 2- and 4-bit weights, not 8-bit. sha256 1bccf05b4bfcfdf1bfebb6e6dd43ddcd0ee20c534b46ecfb4d07147d62f0c82c.

It is not the same model to the last decimal. Over 720 questions (the reference package's five preset question sets on 30 texts), compared with the PyTorch fp32 model:

Build Size Same answer Answers changed Median / 95th pct / max probability shift
laya_fp16.onnx 900 MB 100.0% 0 0.0002 / 0.0014 / 0.009
laya_q8.onnx 633 MB 99.3% 5 (2 yes/no) 0.0016 / 0.0147 / 0.082

A 4-bit build was measured too (449 MB, 92.6% same answer) and is not published.

Reproduce

git clone https://github.com/wexare-ai/browser-laya && cd browser-laya
uv venv -p 3.12 tools/.venv
uv pip install -p tools/.venv/bin/python laya onnx onnxscript onnxruntime
tools/.venv/bin/python tools/export_onnx.py

The script exports with torch.onnx.export (dynamo=True), checks each stage against PyTorch, and stops if a decision changes.

License and attribution

Apache 2.0, the licence of the original weights; see LICENSE.

  • Model, weights, tokenizer and calibration: Convai Innovations, convaiinnovations/laya.
  • Changes in this repository: converted from PyTorch safetensors to ONNX, with the encoder cast to fp16 as described above. The tokenizer and configuration files are unchanged copies.

weXare is not affiliated with Convai Innovations, and this export is not endorsed by them. "Laya" refers to Convai Innovations' model. Questions about the model itself belong with them; issues with this export go to wexare-ai/browser-laya.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for wexare/laya-onnx

Quantized
(62)
this model