OSTQuant Qwen3-4B W4A4KV4 (INT4-packed, compact)

Bit-packed to INT4 from the raw OSTQuant W4A4KV4 checkpoint. Cannot be loaded via vanilla AutoModelForCausalLM.from_pretrained(...) — the OSTQuant runtime is required to unpack and evaluate correctly.

  • Compressed from ~8.9 GB (raw .bin) to ~2.9 GB (3× compression).
  • Weights on 4-bit grid. Activations AND KV cache also quantized to 4 bits at inference.
  • Reference PPL when loaded via OSTQuant runtime: 17.02 (WikiText-2 seq 2048, raw .bin).

Source

  • Raw: TrojAI/ostquant_qwen3_4b_w4a4kv4
  • Packing script: rebuild_int4_from_raw.py
Downloads last month
-
Safetensors
Model size
1.0B params
Tensor type
I32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TrojAI/ostquant_qwen3_4b_w4a4kv4_int4_v2

Finetuned
Qwen/Qwen3-4B
Finetuned
(1062)
this model