1qh/chandra-ocr-2-oq2-mlx

datalab-to/chandra-ocr-2 quantized with oQ (oMLX v0.5.3) at level 2.0, data-driven mixed precision: oQ measures each layer's quantization error through calibration and allocates bits where the measurement says they matter, rather than by a fixed per-tensor rule.

Base datalab-to/chandra-ocr-2
oQ level 2.0 (effective ~2.8–3.0 bits/weight)
Size 2.3 GB (from 9.1 GB bf16)
Sensitivity entries 248
oMLX v0.5.3

Table-column fidelity: PARTIAL / unstable. On the OCR fidelity check this build produced a correct table on one page and an unreadable one on another, with no column shift observed. Table output is unreliable at this level; prefer 2.5 or 5+ when structure matters.

Vision weights are kept fp16; only the language tower is quantized.

Use

from mlx_vlm import load, generate
model, processor = load("1qh/chandra-ocr-2-oq2-mlx")
Downloads last month
24
Safetensors
Model size
0.8B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1qh/chandra-ocr-2-oq2-mlx

Quantized
(34)
this model