Instructions to use RESMP-DEV/LFM2.5-Encoder-350M-Code-MXFP8-GPTQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use RESMP-DEV/LFM2.5-Encoder-350M-Code-MXFP8-GPTQ with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir LFM2.5-Encoder-350M-Code-MXFP8-GPTQ RESMP-DEV/LFM2.5-Encoder-350M-Code-MXFP8-GPTQ
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
File size: 4,827 Bytes
41aeb21 877c723 41aeb21 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | ---
license: other
license_name: lfm1.0
license_link: LICENSE
base_model: LiquidAI/LFM2.5-Encoder-350M
pipeline_tag: feature-extraction
library_name: mlx
tags:
- code
- embeddings
- feature-extraction
- mlx
- mxfp8
- gptq
- quantized
---
# LFM2.5 Encoder 350M Code MXFP8-GPTQ
This is a modified RESMP.DEV research release derived from
[`LiquidAI/LFM2.5-Encoder-350M`](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) at revision `b886781f7c6f10ca9b7096e21b83e30a073c2f39`. It is
not an official Liquid AI release. We removed the masked-language-model head and
contrastively fine-tuned the full bidirectional encoder for multilingual code retrieval.
## Quantization finding
This is a research artifact, not an automatic recommendation to replace the BF16 model.
Activation calibration is compared with matched native round-to-nearest quantization and
the complete machine-readable receipts are included so mobile and Apple-Silicon users
can evaluate the size, latency, memory, and quality tradeoff themselves.
## Held-out retrieval results
All rows use the same untouched 6,995-pair multilingual test set, 1,200-character query
and 4,000-character passage caps, query token cap 512, and passage token cap 2,048.
Higher is better. RTN is a matched quantization control; Nomic and Jina are external
service baselines, not architecture-matched controls.
| Model | MRR | R@1 | R@5 | R@10 | NDCG@10 | Python MRR | TypeScript MRR | Artifact |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| LFM2.5 350M BF16 | 0.3705 | 0.2996 | 0.4422 | 0.5061 | 0.3963 | 0.7969 | 0.1917 | 713.7 MB |
| LFM2.5 350M calibrated MXFP4 | 0.1585 | 0.1169 | 0.1971 | 0.2317 | 0.1697 | 0.5795 | 0.0542 | 291.8 MB |
| LFM2.5 350M RTN MXFP4 | 0.0555 | 0.0422 | 0.0618 | 0.0773 | 0.0576 | 0.3204 | 0.0157 | 291.8 MB |
| LFM2.5 350M calibrated MXFP8 | 0.3710 | 0.3019 | 0.4430 | 0.5045 | 0.3962 | 0.7985 | 0.1903 | 435.5 MB |
| LFM2.5 350M RTN MXFP8 | 0.3684 | 0.2965 | 0.4427 | 0.5054 | 0.3945 | 0.7957 | 0.1932 | 435.4 MB |
| Nomic v1.5 service | 0.5439 | 0.4968 | 0.5954 | 0.6236 | 0.5595 | 0.9289 | 0.3617 | service |
| Jina calibrated MXFP4 | 0.6645 | 0.6133 | 0.7221 | 0.7571 | 0.6832 | 0.9462 | 0.5057 | 1167.7 MB |
A separate BF16 cross-runtime run on `NVIDIA GeForce RTX 3090 Ti` with PyTorch `2.13.0+cu130` produced MRR 0.3709, 616.2 queries/s, 139.1 passages/s, and 1109.0 MB peak CUDA allocation. CUDA throughput is reported separately and is not compared directly with Metal.
A paired 10,000-sample bootstrap estimates calibrated MXFP8 minus BF16 MRR at +0.0005, with a 95% interval of [-0.0010, +0.0021]. A point estimate whose interval crosses zero is not presented as a
quality win.
## Usage
```bash
git clone https://github.com/RESMP-DEV/calibrated-code-embeddings
cd calibrated-code-embeddings
uv sync --extra mlx
CODE_EMBEDDING_MODEL_PATH=/path/to/this-model code-embedding-serve --port 1235
```
The service exposes `POST /v1/embeddings`. It runs the bidirectional LFM2.5 body
directly with MLX; LM Studio is not required. Prefix retrieval queries with `query: `
and candidate code with `passage: ` when calling the model directly.
## Training and data receipts
Full-backbone symmetric in-batch InfoNCE training used 24,626
language-balanced pairs selected from the 42,626-row source training
split, two epochs, batch size 32, learning rate 2e-5, temperature 0.05, and seed 17. The
training report records the NVIDIA RTX A6000 runtime and validation history.
- `train`: 42,626 rows, SHA-256 `426ebfaad34b14d7627ba6e668ae36e08e548c9d057b0edc208bcfa6fe527629`
- `validation`: 5,319 rows, SHA-256 `9ac88b3138de4ca94c2ef3a87ccf19381fc76265c2bc9d65b4791983d0315096`
- `test`: 6,995 rows, SHA-256 `9ed10842a12132b6bfb5421df1e2f88dbcfbf6f6e960f36b22eb9ea6e3c72315`
- `calibration`: 4,096 rows, SHA-256 `ee9edaf80a6854c18053b96521090a51bdb76642abeb98618d7aed36e70b6de9`
The corpus combines pinned CodeSearchNet data with pinned permissively licensed code
repositories. Exact and token 8-gram near-duplicates were removed with test-before-
validation-before-train precedence. See `corpus_receipt.json`, `source_receipt.json`,
`training_report.json`, `quantization_report.json` when present, `benchmarks/`, and
`artifact_manifest.json` for machine-readable evidence.
## License and attribution
The weights retain the LFM Open License v1.0 in `LICENSE`, including its attribution and
commercial-use conditions. `MODIFICATIONS.md` identifies RESMP.DEV's changes. The
[training and quantization workbench](https://github.com/RESMP-DEV/calibrated-code-embeddings)
is separately MIT licensed.
## Citation
```bibtex
@article{liquidAI2026Encoders,
author = {Liquid AI},
title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-encoders},
}
```
|