File size: 2,664 Bytes
672e862
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
947bb7c
 
672e862
947bb7c
672e862
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
947bb7c
672e862
 
 
 
 
 
 
 
947bb7c
672e862
 
947bb7c
672e862
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
---
license: apache-2.0
base_model: Qwen/Qwen3-8B
tags:
- quantization
- trellis-coded-quantization
- qtip
- webgpu
- 3-bit
- on-device
library_name: transformers
---

# Qwen3-8B — 3-bit trellis-quantized for WebGPU

A **3-bit trellis-coded quantization (TCQ)** of [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B),
packed to run its **full forward pass in a web browser on WebGPU** — no CUDA, no server.
Runs on a **4 GB GPU**: verified live in Chrome on a laptop RTX 3050 Ti, where it completed
*"The capital of France is"**"Paris. The capital of Italy is Rome. The capital of Germany is Berlin."*

This is part of **[Trellis WebGPU](https://github.com/Azimml/trellis-webgpu)** —
the first trellis-coded quantization decoder (the QTIP / EXL3 quality tier) to run
outside CUDA. See the repo for the quantizer, the WGSL kernels, and the full
verification harness.

## Quality (WikiText-2 perplexity, no fine-tuning)

| | fp16 | **TCQ 3-bit (this model)** | vs fp16 |
|---|---|---|---|
| Qwen3-8B | 9.02 | **10.34** | 1.15× |

Trellis-coded quantization is the current quality frontier for low-bit LLM weights,
beating scalar quantization (GPTQ / AWQ / GGUF) at equal bitrate. The K=2 quantizer in
this project reproduces the QTIP paper's published rate–distortion (MSE 0.0739 vs 0.0733).

## This packed model

- **36 layers**, packed to **~3 bits/weight** (2.9 GB on disk).
- runs in a browser on a **4 GB GPU** — verified live in Chrome on a laptop **RTX 3050 Ti (4 GB)**, generating the completion above without running out of memory (256-token KV cache; `maxSeq` is configurable for longer contexts).
- Verified: the full model runs through the exact shipping WGSL shaders and generates
  coherent, factually-correct text; kernels match NumPy to `<1e-6`; 135M logits are
  top-5 exact vs PyTorch.

## Format

This is **not** a standard `transformers` checkpoint. It is a packed 3-bit format
(`manifest.json` + sharded `.bin` weights + IP metadata) designed for the WebGPU runtime
in the [Trellis WebGPU](https://github.com/Azimml/trellis-webgpu) repo. To run it:

```bash
git clone https://github.com/Azimml/trellis-webgpu
# place these files under web/model_8b/ , then:
cd web && python3 -m http.server 8000   # open index.html in a WebGPU browser
# or verify headlessly on your GPU:
python scripts/run_packed_headless.py web/model_8b "The capital of France is" 40
```

## License & credit

Quantized derivative of [Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) (Apache-2.0),
distributed under the base model's license. Method: QTIP (Tseng et al., NeurIPS 2024)
and EXL3 (turboderp). Quantization + WebGPU port: independent reimplementation.