Feature Extraction
MLX
Safetensors
qwen3
embeddings
sentence-similarity
quantization
omlx
q4
4-bit precision
Instructions to use TiGa-RCE/Qwen3-Embedding-8B-MLX-Q4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TiGa-RCE/Qwen3-Embedding-8B-MLX-Q4 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3-Embedding-8B-MLX-Q4 TiGa-RCE/Qwen3-Embedding-8B-MLX-Q4
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
File size: 2,921 Bytes
8afd760 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 | ---
license: apache-2.0
library_name: mlx
pipeline_tag: feature-extraction
base_model: TiGa-RCE/Qwen3-Embedding-8B-MLX-BF16
tags:
- mlx
- embeddings
- feature-extraction
- sentence-similarity
- quantization
- omlx
- q4
- 4-bit
---
# Qwen3-Embedding-8B — MLX Q4
> **Capacity-first experimental checkpoint:** every Q4-family build in this sweep failed the predeclared minimum aligned-cosine fidelity gate. It is published as an experimental comparator, not as a BF16-equivalent embedding model.
This is the **Q uniform affine quantization** checkpoint from a matched local embedding-quantization experiment. It is published with explicit lineage, calibration evidence where applicable, and the bounded evaluation result that accompanied the conversion.
## Provenance and lineage
- Upstream model: [`Qwen/Qwen3-Embedding-8B`](https://huggingface.co/Qwen/Qwen3-Embedding-8B)
- Upstream revision recorded for publication: `1d8ad4ca9b3dd8059ad90a75d4983776a23d44af`
- Revision evidence: exact source snapshot retained in the local Hugging Face download metadata
- Direct parent: [`TiGa-RCE/Qwen3-Embedding-8B-MLX-BF16`](https://huggingface.co/TiGa-RCE/Qwen3-Embedding-8B-MLX-BF16)
- Conversion rule: every quantized checkpoint branches directly from the family MLX BF16 checkpoint; no lossy checkpoint was used to create another.
- Quantization: Q uniform affine quantization, nominal 4-bit, group size 64
- Local conversion stack: oMLX 0.5.3, mlx-lm 0.31.3, MLX 0.32.0
- Full collection: [MLX Embedding Quantization Matrix](https://huggingface.co/collections/TiGa-RCE/mlx-embedding-quantization-matrix-q-oq-oqe-at-4-6-8-bit-6a68d11afb238d4fe967d70b)
`PROVENANCE.json` contains machine-readable lineage and SHA-256 hashes for the published weight files. No importance matrix was used for this checkpoint.
## Bounded local evaluation
| Metric | Result |
|---|---:|
| Top-1 retrieval | 1.000 |
| MRR | 1.000 |
| Mean aligned cosine vs BF16 | 0.973145 |
| Minimum aligned cosine vs BF16 | 0.960922 |
| Score RMSE vs BF16 | 0.017052 |
| Queries with rank change | 0 |
| Predeclared gate | FAIL |
The evaluation used 24 frozen query/document pairs, the upstream query instruction recipe, last-token pooling, L2 normalization, and direct comparison with vectors from the family BF16 checkpoint. This is an engineering smoke test, not MTEB and not a claim of universal quality. Retrieval success and representation fidelity are reported separately.
## Runtime scope
This checkpoint targets Apple Silicon through MLX/oMLX. CUDA and PyTorch results are a separate control lane and must not be interpreted as measurements of MLX/Metal kernel performance.
## License and attribution
Apache-2.0, following the upstream model card. The original model authors retain attribution for the upstream model; this repository contains a local MLX conversion or quantized derivative prepared by TiGa-RCE for reproducibility research.
|