Text Generation
Transformers
Safetensors
nvfp4
compressed-tensors
vllm
blackwell
conversational
custom_code
8-bit precision
Instructions to use AtomicChat/Ling-3.0-flash-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AtomicChat/Ling-3.0-flash-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AtomicChat/Ling-3.0-flash-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AtomicChat/Ling-3.0-flash-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AtomicChat/Ling-3.0-flash-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtomicChat/Ling-3.0-flash-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Ling-3.0-flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AtomicChat/Ling-3.0-flash-NVFP4
- SGLang
How to use AtomicChat/Ling-3.0-flash-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AtomicChat/Ling-3.0-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Ling-3.0-flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AtomicChat/Ling-3.0-flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/Ling-3.0-flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AtomicChat/Ling-3.0-flash-NVFP4 with Docker Model Runner:
docker model run hf.co/AtomicChat/Ling-3.0-flash-NVFP4
File size: 5,543 Bytes
91e518c 2ebb583 91e518c 2ebb583 91e518c 2ebb583 91e518c 2ebb583 91e518c ba1a22a 91e518c f3c3fe6 91e518c ba1a22a f3c3fe6 ba1a22a 91e518c 2ebb583 91e518c 2ebb583 91e518c ba1a22a 2ebb583 9b8ad9c f7bcaea 9b8ad9c 2ebb583 f7bcaea 2ebb583 9b8ad9c f7bcaea 9b8ad9c 2ebb583 9b8ad9c f3c3fe6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 | ---
license: mit
base_model: inclusionAI/Ling-3.0-flash
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- nvfp4
- compressed-tensors
- vllm
- blackwell
library_name: transformers
---
# Ling-3.0-flash — NVFP4 (safetensors)
[`inclusionAI/Ling-3.0-flash`](https://huggingface.co/inclusionAI/Ling-3.0-flash) quantized to NVFP4 in the `compressed-tensors` format, for vLLM and TensorRT-LLM.
**76.9 GB across 16 shards.** The GGUF builds for llama.cpp live in
[`AtomicChat/Ling-3.0-flash-NVFP4-GGUF`](https://huggingface.co/AtomicChat/Ling-3.0-flash-NVFP4-GGUF).
**Mean KL divergence against the bf16 baseline: 0.0409, top-1 agreement 95.9%** - more accurate than
our GGUF NVFP4 at a comparable size. Method and full numbers below.
## What this is
Scheme `NVFP4A16`, as written into `config.json`:
```
format: nvfp4-pack-quantized | producer: compressed-tensors
group_0: bits=4 type=float strategy=tensor_group group=16 symmetric=True
targets: ['Linear'] activations: None
```
4-bit E2M1 weights, one UE4M3 scale per 16 weights, **plus a per-tensor global scale** — `tensor_group`
is exactly the part the GGUF block layout has no room for, and it does put this build closer to the
original than our GGUF NVFP4 at the same bit width. Activations stay bf16, so no calibration data is
involved.
Same layout policy as the rest of our line: only the routed experts, 97.1% of the weights, are quantized. Everything cheap in parameters but expensive in error is left alone:
```python
QuantizationModifier(
targets=["Linear"],
scheme="NVFP4A16",
ignore=[
"lm_head",
r"re:.*mlp\.gate$", # the router: an error changes which expert runs
r"re:.*shared_experts\..*", # seen by every token
r"re:.*attention\..*", # MLA, KDA projections and gates
],
)
```
In the written config those patterns are expanded into an explicit list of more than 400 modules:
every `attention.*` projection (the KDA `f_proj`, `b_proj`, `g_proj` and the MLA `kv_a_proj_with_mqa`,
`kv_b_proj`, `dense`), every `shared_experts.*` and `lm_head`. Only the routed experts are in NVFP4.
The MTP block (layer 42) is kept here in bf16. The GGUF converter drops it, so if multi-token
prediction lands in vLLM, this build has what it needs and the GGUF one does not.
## Running it
Ling 3.0 is not in upstream vLLM yet, but the support is open as
[PR #51045](https://github.com/vllm-project/vllm/pull/51045) and can be installed straight from it,
no need to wait for the merge:
```bash
git clone https://github.com/vllm-project/vllm && cd vllm
gh pr checkout 51045 # or: git fetch origin pull/51045/head:ling3 && git checkout ling3
VLLM_USE_PRECOMPILED=1 uv pip install --editable . --torch-backend=auto
vllm serve AtomicChat/Ling-3.0-flash-NVFP4 --trust-remote-code --tensor-parallel-size 4
```
The model authors also publish a fork, [`inclusionAI/vllm-ling-v3`](https://github.com/inclusionAI/vllm-ling-v3),
which works as well; the PR is closer to what will eventually ship upstream.
On arm64 machines such as DGX Spark, `VLLM_USE_PRECOMPILED=1` will not help - those wheels are built
for x86_64 - so expect a full source build there.
NVFP4 has native tensor-core support on Blackwell (sm_100, sm_120, sm_121). Elsewhere it runs through
dequantization, without the speedup.
## Measuring it
`NVFP4A16` is weight-only: activations stay bf16, so a model whose weights have been round-tripped
through the FP4 grid is mathematically the same model as this compressed checkpoint. That makes the
measurement easy to keep honest - the checkpoint is expanded back to dense bf16 and then goes through
**the exact same harness as every other rung**: `llama-perplexity --kl-divergence` against the same
bf16 baseline, the same 92 chunks of held-out text at 4096 tokens, the same 4x RTX PRO 6000 machine.
No second definition of the metric, no cross-runtime fudge factor.
The expansion was checked against `compressed_tensors`' own `unpack_fp4_from_uint8` before use
(bit-identical) and against the original bf16 weights (relative L2 error 0.095 on the first tensor,
which is what FP4 costs).
| | size | mean KLD | top-1 | RMS dp | PPL |
| --- | --- | --- | --- | --- | --- |
| bf16 baseline | 248.9 GB | 0 | 100% | 0% | 4.2059 |
| **NVFP4 (this repo)** | **76.9 GB** | **0.0409 +- 0.0010** | **95.93%** | **6.21%** | 4.2255 |
| GGUF `AD-IQ4_XXS` | 69.3 GB | 0.0329 | 96.44% | - | - |
| GGUF `AD-NVFP4` | 72.3 GB | 0.0536 | 94.87% | - | - |
Distribution of the divergence: median 0.0015, 99th percentile 0.448, maximum 21.9. As with every
4-bit build, the mean is carried by a small tail of tokens where the router flips to a different
expert, not by a uniform blur.
Two things separate this build from the GGUF NVFP4 of the same width, and they pull in opposite
directions. In its favour, `tensor_group` carries a per-tensor global scale that the GGUF block has no
room for, and the ignore list keeps attention and the shared experts in bf16 rather than Q8_0. Against
it, the two dense FFN layers at the front of the model land in NVFP4 here, where the GGUF build keeps
them at Q8_0. On net it comes out 0.013 ahead.
Worth stating plainly: at 4-bit, NVFP4 is not the accuracy-per-byte winner in this line. `AD-IQ4_XXS`
is both smaller and closer to the original. NVFP4 is here for the Blackwell tensor cores, not for the
error curve.
Raw logs and json: [`AtomicChat/Ling-3.0-flash-GGUF-metrics`](https://huggingface.co/datasets/AtomicChat/Ling-3.0-flash-GGUF-metrics).
|