clef-flash-oQ4e

This model was quantized using oQ mixed-precision quantization.

A mixed-precision MLX checkpoint of Cloudflare/clef-flash, a decision model that answers typed questions (noul, choice, score) about text and images with a probability for each option. It includes the Qwen3.5 language backbone, the vision encoder and the joint schema head.

The joint schema head (joint_head.safetensors, joint_head_config.json) is copied unchanged from the source. oMLX serves this checkpoint through POST /v1/systemone once decision model support (jundot/omlx#4315) is available.

The effective average is 5.574 bits/weight for the complete checkpoint.

Quantization and Bit Distribution

  • 4-bit affine quantization, group size 64, for most language-backbone projections and the token embeddings.
  • 5-bit affine quantization, group size 64, for projections boosted by measured layer sensitivity.
  • BF16 for the vision encoder, the joint schema head, norms and the remaining small tensors.

The oQ4e allocation uses measured layer sensitivity and importance-matrix calibration (128 samples x 512 tokens, built-in code and multilingual set).

Storage format Logical weights Share of total Tensor storage Effective bits/weight
Affine 4-bit, group 64 8.288B 84.469% 4.342 GiB 4.50
Affine 5-bit, group 64 0.665B 6.780% 0.426 GiB 5.50
BF16 0.859B 8.751% 1.599 GiB 16.00
Total 9.811B 100% 6.367 GiB 5.574

Storage Breakdown

Component Logical weights Storage
Language backbone 9.234B 5.681 GB / 5.291 GiB
Vision encoder 0.456B 0.912 GB / 0.849 GiB
Joint schema head 0.122B 0.244 GB / 0.227 GiB
Total 9.811B 6.837 GB / 6.367 GiB

License

Apache-2.0, following the source model. See LICENSE.

Downloads last month
57
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jundot/clef-flash-oQ4e

Finetuned
Qwen/Qwen3.5-9B
Quantized
(35)
this model