🚀 Qwen3.8-Flash-Coder (Selective INT8 Quantized - ~19.7 GB)

Qwen3.8-Flash-Coder (Selective INT8) is an ultra-compressed, zero-loss MoE architecture derived from Qwen/Qwen3.8-Flash-Next (335GB).

🌟 Architectural Breakthrough: Selective Quantization

  • 100% Native BF16 Precision for Router & Attention: The routing gate matrices across all 48 layers are preserved in full BF16 precision, eliminating 100% of quantization noise and softmax distortion.
  • Selective INT8 Compression for 128 FFN Experts: Compresses 93% of the model parameters into INT8 per-channel quantization, reducing total memory from 65.32 GB down to ~19.72 GB.
  • Fits on a Single 32GB GPU: Seamlessly deployable on 1x NVIDIA RTX 5000 Ada (32GB), RTX 4090 (24GB), or A100/H100 with ample headroom for KV Cache.

📊 Technical Specifications

  • Base Architecture: 48 MoE Layers (128 Routed Experts / Layer, Top-8 Active)
  • Model Size: ~19.7 GB Safetensors (Single GPU Native Deployment)
  • Accuracy Pass Rate: 100.0% (4/4 Pass) on Python & Rust standard evaluation benchmarks.
  • License: Apache 2.0
Downloads last month
18
Safetensors
Model size
35B params
Tensor type
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jab1718/qwen3.8-flash-coder-selective-int8

Finetuned
(24)
this model