Onyx-mE5-Small (4-bit MLX)

4-bit MLX quantization of intfloat/multilingual-e5-small (MIT), for the Onyx on-device RAG embedder (OnyxEmbed). group_size 64, bits 4 on the attention/FFN dense matrices + word_embeddings; LayerNorms, biases, and positional embeddings stay fp. ~71 MB (from ~449 MB fp32).

Usage keeps e5 conventions: mean pooling, and the query: / passage: prefixes. Verified recall-neutral vs fp32 (English + cross-lingual). Inherits the base model MIT license.

Downloads last month
152
Safetensors
Model size
18.6M params
Tensor type
F32
I64
U32
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for wabibito/Onyx-mE5-Small

Finetuned
(185)
this model