EXL3 quant of apple/LensVLM-9B, a vision-language model built on Qwen/Qwen3.5-9B. Requires ExLlamaV3 1.5.1.
Model derivative and attribution
This repository is an unofficial Model Derivative of the Apple Machine Learning Research Model apple/LensVLM-9B, and is identified as such. It is not produced, reviewed, or endorsed by Apple.
"Apple Machine Learning Research Model is licensed under the Apple Machine Learning Research Model License Agreement." The Agreement is included here as LICENSE, and it limits this derivative, like the source model, to Research Purposes.
Changes made to the source model:
- Decoder, output-head and vision-tower linear weights are replaced by EXL3 trellis quantizations (
quant_format: exl3, codebookmul1). Embeddings, every norm, the patch-embed convolution and the position embedding stay at source precision. text_config.mtp_num_hidden_layerschanged from 1 to 0. The sourceconfig.jsondeclares one MTP layer, but this checkpoint'smodel.safetensorscarries nomtp.*tensors, so the field is false for these weights. Nothing was added and nothing was retrained.preprocessor_config.jsonadded, copied unchanged from this repository'sprocessor_config.jsonimage_processorobject, which is the only image-preprocessor config the source ships.- Apple's page-scan and selective-expansion tool loop is not part of the weights and is not included here.
Quants
4.00 bits per weight / H6 / V6
5.00 bits per weight / H6 / V6
6.00 bits per weight / H6 / V6
3.00 bits per weight / H6 / V6
How these quants were built
The bit map follows turboderp's published Qwen3.5-9B self-calibrated pack, with the output head at 6 bits, rather than the Qwen3.8-27B head schedule, because this is a 9B model. The architecture is dense. There is no n-gram table, there is no MTP block to quantize, and -hq does not raise any layer. Vision bits are passed explicitly, because older quant configs omit them.
4.00 bpw:
-b 4 -hb 6 -vb 6. Decoder layers are 4 bits. The output head is 6 bits. This model has no MTP block. The vision tower is 6 bits.5.00 bpw:
-b 5 -hb 6 -vb 6. Decoder layers are 5 bits. The output head is 6 bits. This model has no MTP block. The vision tower is 6 bits.6.00 bpw:
-b 6 -hb 6 -vb 6. Decoder layers are 6 bits. The output head is 6 bits. This model has no MTP block. The vision tower is 6 bits.3.00 bpw:
-b 3 -hb 6 -vb 6. Decoder layers are 3 bits. The output head is 6 bits. This model has no MTP block. The vision tower is 6 bits.
The calibration text is not the converter default. An 8 bpw pack of the same BF16 weights was built first, and sc_trace.py sampled 250 rows by 2048 columns from it with enable_thinking set; a short trace stays short rather than being padded, and each pack records the row count it actually calibrated on in its quantization_config.json. -cd only replaces the text Hessian corpus, so the vision tower is quantized without calibration. Every pack listed here is a fresh conversion of the BF16 weights with -cd pointed at that trace — none is a requant of the 8 bpw pack, and none was built with sc_measure or sc_optimize.
Source weights: apple/LensVLM-9B.
- Downloads last month
- 34