Instructions to use wexare/laya-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use wexare/laya-onnx with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Laya 路 ONNX (fp16) for the browser
An unofficial ONNX export of convaiinnovations/laya,
the open-source "System One" decision model by Convai Innovations, made to run client-side with
ONNX Runtime Web on WebGPU or WebAssembly. Laya takes a state and typed questions and scores every
option in one forward pass. It returns calibrated probabilities, not generated text.
This is the default model of @wexare/laya-web
(source), which loads it automatically:
import { LayaWorkerClient } from "@wexare/laya-web";
Files
| File | What it is |
|---|---|
laya_fp16.onnx |
The model: one self-contained file, 900,153,849 bytes, opset 18 |
tokenizer/tokenizer.json, tokenizer/tokenizer_config.json |
The checkpoint's ModernBERT tokenizer, mirrored unchanged |
rl_agent_config.json |
Calibration temperatures and length limits, mirrored unchanged |
export.json |
Export metadata, including the model's sha256 |
laya_q8.onnx |
Experimental 8-bit build, WebAssembly only: 632,942,133 bytes, see below |
laya_q8.json |
Its metadata and measured agreement |
sha256 of laya_fp16.onnx: b2125c7a1044e78807ee718ff4e65c9892b6b55aeeeee64f31fffffd76871b0a
Graph
inputs input_ids [B,L] int64, attention_mask [B,L] int64,
marker_pos [B,K] int64, marker_mask [B,K] bool, qtype [B] int64
outputs logits [B,K] float32, act_logits [B,2] float32
Batch, sequence and option count are all dynamic. Masked option slots return -1e4. Turning
logits into answers (per-option-count temperature, softmax, confidence) follows the reference
laya package and is done by the library, not the graph.
Precision
Mixed precision, set up in PyTorch before export:
- The ModernBERT-large encoder (about 95% of the weights) is in fp16, with every LayerNorm computing in fp32.
- The decision head (2 transformer layers, option scorer, act head) stays in fp32.
- Inputs and outputs are the types listed above; no fp16 tensors cross the graph boundary.
Verification
Against the reference laya 0.3.3 package, on 9 fixture cases covering single and multi-question
batches, 2 to 14 options, a 600-token truncated state, structured and non-ASCII inputs:
- fp32 export vs PyTorch: identical decisions, maximum probability difference 0.00000.
- This fp16 file vs PyTorch fp32 (ONNX Runtime 1.30, CPU): identical decisions, maximum probability difference 0.00068.
- Through
@wexare/laya-webon onnxruntime-web 1.30 (WASM): 9/9 end-to-end parity tests pass: same chosen option and same yes/no verdict, with probabilities within 0.02 of the Python model.
Experimental: 8-bit (laya_q8.onnx)
A smaller build for testing on phones, where the fp16 model can exhaust the browser's memory while
loading. Its MatMul weights are quantized to 8 bits with ONNX Runtime's MatMulNBits (block 32,
symmetric); everything else stays fp32. WebAssembly only: onnxruntime-web's WebGPU
MatMulNBits kernel accepts 2- and 4-bit weights, not 8-bit. sha256 1bccf05b4bfcfdf1bfebb6e6dd43ddcd0ee20c534b46ecfb4d07147d62f0c82c.
It is not the same model to the last decimal. Over 720 questions (the reference package's five preset question sets on 30 texts), compared with the PyTorch fp32 model:
| Build | Size | Same answer | Answers changed | Median / 95th pct / max probability shift |
|---|---|---|---|---|
laya_fp16.onnx |
900 MB | 100.0% | 0 | 0.0002 / 0.0014 / 0.009 |
laya_q8.onnx |
633 MB | 99.3% | 5 (2 yes/no) | 0.0016 / 0.0147 / 0.082 |
A 4-bit build was measured too (449 MB, 92.6% same answer) and is not published.
Reproduce
git clone https://github.com/wexare-ai/browser-laya && cd browser-laya
uv venv -p 3.12 tools/.venv
uv pip install -p tools/.venv/bin/python laya onnx onnxscript onnxruntime
tools/.venv/bin/python tools/export_onnx.py
The script exports with torch.onnx.export (dynamo=True), checks each stage against PyTorch,
and stops if a decision changes.
License and attribution
Apache 2.0, the licence of the original weights; see LICENSE.
- Model, weights, tokenizer and calibration: Convai Innovations,
convaiinnovations/laya. - Changes in this repository: converted from PyTorch safetensors to ONNX, with the encoder cast to fp16 as described above. The tokenizer and configuration files are unchanged copies.
weXare is not affiliated with Convai Innovations, and this export is not endorsed by them. "Laya" refers to Convai Innovations' model. Questions about the model itself belong with them; issues with this export go to wexare-ai/browser-laya.
Model tree for wexare/laya-onnx
Base model
convaiinnovations/laya