qwen35-2b-webgpu / README.md
borkiss's picture
Upload folder using huggingface_hub
1399b85 verified
|
Raw
History Blame Contribute Delete
1.08 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.5-2B
tags: [webgpu, wgsl, gptq, int4, browser, text-generation, deltanet]
---
# qwen35-2b-webgpu — GPTQ-int4 weights
Custom GPTQ-int4 quant (group 32, byte-sliced nibbles, single tied q4 copy of
the 248k-vocab embedding serving both the gather and the logits cascade) of
[Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) (text part), in the
wire format of the
[qwen35-2b-webgpu](https://huggingface.co/spaces/borkiss/qwen35-2b-webgpu-demo) browser runtime.
**The first hybrid gated-DeltaNet + gated-attention LLM running in the
browser** — pure WebGPU, hand-written WGSL kernels (chunked WY-representation
delta-rule prefill, fused recurrent decode step, f32 state), no ONNX /
transformers.js / GGUF.
- ~1.07 GB; GPTQ calibrated on **chat-formatted** passages.
- The runtime is verified against a PyTorch f32 reference (per-layer rms,
cascaded argmax bit-exact vs the exact-logits path).
Try it: **https://huggingface.co/spaces/borkiss/qwen35-2b-webgpu-demo** · run /web/bench.html
on your GPU and share the .log.