Qwen3-4B W2G128 Block-AP FP32 pre-round (w2_4)
This repository contains the completed sweep_0.8 Block-AP / Stage 1 model and quantizer state for Qwen3-4B, with 2-bit weights and group size 128. It preserves the continuous FP32 weights, scales, and zero points before final half-precision materialization and quantization. The artifact has 903 FP32 tensors, 36 trained blocks, and 37 safetensors shards.
This is the pre-round starting point. It does not include later offline forward-KL, TIP-GKD, reverse-KL, or other OPD training. It is not a full optimizer checkpoint for resuming Block-AP training.
The weights require a QAT-aware loader that preserves the FP32 weight, scale, and zero-point tensors and enables the intended fake-quantized forward path. The standard Hugging Face AutoModelForCausalLM.from_pretrained path should not be assumed to reproduce the quantized model's behavior directly. See QAT_LATENT_MANIFEST.json for artifact semantics and shard checksums.
Source: Qwen3-4B. Block-AP calibration used sweep_0.8, 4,096 training sequences and 64 validation sequences of length 2,048, batch size 2, two epochs, seed 2, and W2G128 quantization. The formal Stage 1 run, saved FP32 tensors, independent reload, and archive all passed validation.
No downstream reasoning accuracy claim is attached to this upload.
- Downloads last month
- 91