MichaelAnthony commited on
Commit
4e1c80a
·
verified ·
1 Parent(s): f136895

Update README for corrected quantization (527 layers, embeddings quantized)

Browse files
Files changed (1) hide show
  1. README.md +22 -20
README.md CHANGED
@@ -17,9 +17,9 @@ tags:
17
  # Gemma 4 E2B SnowFox MLX 6-bit (affine, group 64)
18
 
19
  Standard MLX-VLM 6-bit affine weight quantization of the SnowFox model —
20
- the MLX equivalent of GGUF `Q6_K`. A genuine MLX-VLM package (quantized
21
- `safetensors` + `config.json` carrying a `quantization` field), not a GGUF
22
- file or a renamed HF checkpoint.
23
 
24
  SnowFox is a language-only LoRA merge based on Google's Gemma 4 E2B
25
  instruction QAT-derived checkpoint. The image and audio towers were frozen
@@ -35,18 +35,20 @@ tokenizer needed by MLX-VLM.
35
 
36
  ## What is quantized
37
 
38
- - **525 linear layers** (`q/k/v/o` projections, MLP gate/up/down, multimodal
39
- projectors) are 6-bit affine quantized: packed `uint8` `weight` (4 values per
40
- 3 bytes) + float16 `scales`/`biases`.
41
- - **Embeddings, layer norms, convolutions, and biases stay float16** — matching
42
- MLX-VLM's standard Linear-only affine quantization. The large per-layer input
43
- embedding (`embed_tokens_per_layer`) is retained at full precision, so the
44
- model needs ~7 GB unified memory on Apple Silicon.
 
45
 
46
  ## Package contents
47
 
48
- - `model-00001-of-00003.safetensors` `model-00003-of-00003.safetensors`: the
49
- one 6-bit MLX model (~7.4 GB total).
 
50
  - `model.safetensors.index.json`: complete shard map.
51
  - `config.json` (with `quantization` + `quantization_config`), `generation_config.json`,
52
  `processor_config.json`, tokenizer files, and `chat_template.jinja`.
@@ -56,14 +58,14 @@ tokenizer needed by MLX-VLM.
56
  The conversion host has no Apple-Silicon MLX runtime, so the quantized package
57
  was structurally validated before upload:
58
 
59
- - 1,951 source tensors mapped with no missing or extra keys; 525 linear layers
60
- quantized; embeddings left dense.
61
- - Quantized weight format matches the MLX 6-bit affine contract: 4 values
62
- packed per 3 `uint8` bytes (value *k* at bits `[6k, 6k+6)`), dequantization
63
- `scale * q + bias`, group 64.
64
- - Round-trip dequantization of sampled layers (language + audio towers) reproduces
65
- the source weights to within 6-bit precision (max absolute error ≤ ~0.004).
66
- - Config matches MLX-VLM `convert --quantize --q-bits 6` output.
67
 
68
  **Apple-Silicon MLX-VLM inference has not been run.** Treat this as a
69
  structurally validated quantization pending a real Apple-Silicon text / image /
 
17
  # Gemma 4 E2B SnowFox MLX 6-bit (affine, group 64)
18
 
19
  Standard MLX-VLM 6-bit affine weight quantization of the SnowFox model —
20
+ the MLX equivalent of GGUF `Q6_K`. This is a genuine MLX-VLM package
21
+ (quantized `safetensors` + `config.json` carrying a `quantization` field),
22
+ not a GGUF file or a renamed HF checkpoint.
23
 
24
  SnowFox is a language-only LoRA merge based on Google's Gemma 4 E2B
25
  instruction QAT-derived checkpoint. The image and audio towers were frozen
 
35
 
36
  ## What is quantized
37
 
38
+ - **527 linear and embedding layers** (`q/k/v/o` projections, MLP gate/up/down,
39
+ multimodal projectors, and the large embeddings) are 6-bit affine quantized:
40
+ packed `uint32` `weight` (4 values per 3 bytes, low bits first) + float16
41
+ `scales`/`biases`.
42
+ - **Layer norms, convolutions, and biases stay float16** — matching MLX-VLM's
43
+ standard affine quantization. The dense per-layer input embedding
44
+ (`embed_tokens_per_layer`) **is** quantized here, so the package stays under
45
+ ~4.2 GB rather than the ~7 GB a dense embedding would force.
46
 
47
  ## Package contents
48
 
49
+ - `model-00001-of-00002.safetensors` (4,000,099,030 bytes) and
50
+ `model-00002-of-00002.safetensors` (166,857,424 bytes): the 6-bit MLX model
51
+ (~4.2 GB total).
52
  - `model.safetensors.index.json`: complete shard map.
53
  - `config.json` (with `quantization` + `quantization_config`), `generation_config.json`,
54
  `processor_config.json`, tokenizer files, and `chat_template.jinja`.
 
58
  The conversion host has no Apple-Silicon MLX runtime, so the quantized package
59
  was structurally validated before upload:
60
 
61
+ - 1,951 source tensors mapped with no missing or extra keys; 527 linear +
62
+ embedding layers quantized.
63
+ - Quantized weight format matches the MLX affine contract: 6-bit values packed
64
+ 4-per-3-bytes (24-bit little-endian word), dequantization `scale * q + bias`,
65
+ group 64.
66
+ - Round-trip dequantization of sampled layers (attention projections + the
67
+ 2.35B-param `embed_tokens_per_layer`) reproduces the source weights to within
68
+ 6-bit precision (max relative error 1.4%).
69
 
70
  **Apple-Silicon MLX-VLM inference has not been run.** Treat this as a
71
  structurally validated quantization pending a real Apple-Silicon text / image /