Spaces:
Running
Running
| title: README | |
| emoji: ๐ง | |
| colorFrom: yellow | |
| colorTo: red | |
| sdk: static | |
| pinned: false | |
| <p align="center"> | |
| <img src="assets/org-banner.png" alt="Primitive โ inference economics for production LLM systems" width="100%"> | |
| </p> | |
| <p align="center"> | |
| <a href="https://primitive.com"><img src="https://img.shields.io/badge/web-primitive.com-F98424?style=flat-square" alt="primitive.com"></a> | |
| <img src="https://img.shields.io/badge/focus-inference%20economics-F98424?style=flat-square" alt="Inference economics"> | |
| <img src="https://img.shields.io/badge/runtimes-vLLM%20ยท%20ONNX-F98424?style=flat-square" alt="vLLM and ONNX"> | |
| <img src="https://img.shields.io/badge/open%20weights-6%20models-3DD13D?style=flat-square" alt="Six open-weight models"> | |
| </p> | |
| --- | |
| ## Serving LLMs well is two problems | |
| **Which model should answer this request?** and **how cheaply can that model | |
| run?** Most teams solve neither โ they pin one frontier model to every route and | |
| pay frontier prices on the easy majority of their traffic. | |
| We build the two pieces that fix that. | |
| <table> | |
| <tr> | |
| <td width="50%" valign="top"> | |
| ### ๐ Route | |
| A small, fast model that reads the prompt **before** you spend anything on it, | |
| then ranks which LLM is actually worth calling. Runs on CPU next to your API | |
| server โ no GPU, no PyTorch, no extra service to operate. | |
| </td> | |
| <td width="50%" valign="top"> | |
| ### ๐๏ธ Compress | |
| Mixed-precision quantization that shrinks a checkpoint and speeds up serving | |
| **without quietly trading away accuracy** โ and that still loads on the GPUs you | |
| already have, not just the newest ones. | |
| </td> | |
| </tr> | |
| </table> | |
| --- | |
| ## Out now: six open-weight quants | |
| <p align="center"> | |
| <img src="assets/lineup.png" alt="Across the six public quants: 2.5ร to 3.4ร smaller than BF16; 1.72ร to 2.68ร BF16 throughput at concurrency 32; accuracy within 0.5 points of BF16 in either direction, inside the ยฑ0.8 run-to-run band โ ties, not wins" width="100%"> | |
| </p> | |
| Weights-only mixed precision, `compressed-tensors`, stock `vllm serve`. Every | |
| figure below is on the model's own card, next to the run that produced it. | |
| * ๐งฎ **[Qwen3.8-27B โ mixed NVFP4/FP8](https://huggingface.co/primitive-ai/Qwen3.8-27B-mixed-NVFP4-FP8)** | |
| ยท 20.7 GiB, **2.5ร smaller than BF16** and 28% smaller than Qwen's own FP8, with | |
| accuracy tied to BF16 (88.5 vs 88.6 โ and the same weights re-scored 87.5 in a | |
| second sitting, which is what a tie looks like). **Loads on an A100**, not just | |
| Blackwell, and the in-checkpoint MTP head survives, so speculative decoding | |
| still works. | |
| * ๐ชถ **[Laguna-XS-2.1 โ mixed NVFP4/MXFP8](https://huggingface.co/primitive-ai/Laguna-XS-2.1-mixed-NVFP4-MXFP8)** | |
| ยท 19.3 GiB at 4.96 bits/weight, **3.2ร smaller than BF16 and smaller than every | |
| official quant** โ while running faster than all of them at 3,268 tok/s. | |
| * ๐ **[Laguna-S-2.1 โ mixed NVFP4/MXFP8](https://huggingface.co/primitive-ai/Laguna-S-2.1-mixed-NVFP4-MXFP8)** | |
| ยท 219 GB of source weights down to 64 GiB at 4.67 bits/weight โ **it serves | |
| from one 96 GiB GPU**, the smallest and fastest of the 4-bit field, with +23% | |
| single-stream decode over the NVFP4 alternative at a 3 GiB smaller footprint. | |
| * ๐๏ธ **[Muse-Glimmer-30B โ mixed NVFP4](https://huggingface.co/primitive-ai/Muse-Glimmer-30B-mixed-NVFP4)** | |
| ยท A vision-language model at 20.4 GiB, **2.40ร BF16 throughput** and 86.0 | |
| against BF16's 86.2. **The vision tower stays BF16** โ a VLM loses its visual | |
| grounding before it loses its words. | |
| * ๐ฆ **[Ornith-1.5-35B-A3B โ mixed NVFP4/FP8](https://huggingface.co/primitive-ai/Ornith-1.5-35B-A3B-mixed-NVFP4-FP8)** | |
| ยท A 35 B MoE vision-language model at 22.6 GiB, **3.0ร smaller than BF16** and | |
| 39% smaller than the official FP8, scoring 91.5 against BF16's 92.0 โ the only | |
| quant of this model we have measured level with BF16. **The official NVFP4 is | |
| faster and 0.75 GiB smaller**; this build trades that for the closest accuracy | |
| to BF16 and a BF16 output head that keeps the Ampere path loadable. | |
| * โก **[Nemotron-3.5-Lightning-30B-A3B โ mixed INT4/INT8](https://huggingface.co/primitive-ai/Nemotron-3.5-Lightning-30B-A3B-mixed-INT4-INT8)** | |
| ยท A MoE+Mamba hybrid at 19.2 GiB and 5.01 bits/weight, **3.3ร smaller than | |
| BF16 at 2.7ร its throughput** โ and **8.6% smaller and ~7.7% faster than | |
| nvidia's own NVFP4**, from weights-only round-to-nearest with no calibration | |
| data. We make no accuracy claim here: identical weights scored 88.0 and 87.4 | |
| on two runs, so that column is a tie. | |
| --- | |
| ## Still in private beta | |
| The router is not published yet. We're working with a small number of teams | |
| running real production traffic โ the ones who can tell us where the sharp edges | |
| are before everyone else finds them. | |
| **A good fit if you're:** paying per-token to a frontier API on traffic you | |
| suspect doesn't need it, serving an open-weights model at meaningful volume, or | |
| squeezing a large checkpoint onto GPUs you already own. | |
| <p align="center"> | |
| <a href="https://primitive.com"><b>Request access at primitive.com โ</b></a> | |
| </p> | |
| --- | |
| ## How we publish | |
| * **Numbers or nothing.** Every figure on our model cards comes from a held-out | |
| suite and is reproducible from files in the repo. Where a metric is relative, | |
| we say what it's relative to. | |
| * **We document what breaks.** Our cards carry the failure modes, not just the | |
| wins โ the dtype choices that matter, the upstream quirks that bite, and the | |
| optimisations we tried and *rejected*. | |
| * **No silent quality loss.** A compression or routing win that changes answers | |
| isn't a win. When a cheaper option moved outputs too much, we don't ship it โ | |
| and we say so on the card. | |
| * **A tie is a tie.** Our accuracy deltas sit inside the run-to-run spread of the | |
| suite, and we label them as ties rather than dressing them up as wins. | |
| * **Runs on hardware you have.** Older-generation GPU fallbacks for the quants, | |
| CPU-only for the router. Nothing here assumes you just bought the newest rack. | |
| --- | |
| ## Repositories | |
| | repo | what it does | status | | |
| |---|---|---| | |
| | `Qwen3.8-27B-mixed-NVFP4-FP8` | Mixed-precision quant of Qwen3.8-27B. vLLM-native, A100 fallback, MTP head preserved for speculative decoding. | **Open weights** | | |
| | `Laguna-XS-2.1-mixed-NVFP4-MXFP8` | Mixed-precision quant of poolside's Laguna-XS-2.1. Smallest and fastest build of it we've measured. | **Open weights** | | |
| | `Laguna-S-2.1-mixed-NVFP4-MXFP8` | Mixed-precision quant of poolside's Laguna-S-2.1. 219 GB of weights onto a single 96 GiB card. | **Open weights** | | |
| | `Muse-Glimmer-30B-mixed-NVFP4` | NVFP4 quant of the Muse-Glimmer-30B VLM, GPTQ-calibrated. Vision tower left at BF16. | **Open weights** | | |
| | `Ornith-1.5-35B-A3B-mixed-NVFP4-FP8` | Mixed-precision quant of Ornith-1.5-35B-A3B, a 35 B MoE VLM. Vision tower and MTP block left at BF16. | **Open weights** | | |
| | `Nemotron-3.5-Lightning-30B-A3B-mixed-INT4-INT8` | INT4/INT8 quant of nvidia's MoE+Mamba hybrid. Smaller and faster than the official NVFP4, no calibration data. | **Open weights** | | |
| | `llm-router-lite-80m-onnx` | CPU-only router: reads a chat `messages` array, ranks which LLM to call. ONNX, no torch or transformers. | Private beta | | |
| Published weights carry their upstream licence โ Apache-2.0 for the Qwen and | |
| Muse-Glimmer builds, OpenMDW-1.1 for the Laguna and Nemotron builds, MIT for the | |
| Ornith build. Derivative models carry upstream attribution and a `NOTICE` or | |
| `LICENSE` file. | |
| --- | |
| <h3 align="center">Cutting your inference bill?</h3> | |
| <p align="center"> | |
| We'd like to hear what you're serving and where it hurts.<br> | |
| <a href="https://primitive.com"><b>primitive.com</b></a> | |
| </p> | |
| <p align="center"> | |
| <br> | |
| <img src="assets/primitive-logo.png" alt="Primitive" width="34"><br> | |
| <sub><a href="https://primitive.com"><b>primitive</b></a> ยท inference economics for production LLM systems</sub> | |
| </p> | |