File size: 2,239 Bytes
d8c7374
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
---
base_model: Qwen/Qwen3-Reranker-4B
base_model_relation: quantized
library_name: transformers
tags:
- fp8
- compressed-tensors
- llm-compressor
- vllm
- reranker
---

# Qwen3-Reranker-4B-FP8-Dynamic

[Qwen/Qwen3-Reranker-4B](https://huggingface.co/Qwen/Qwen3-Reranker-4B) quantised to FP8 with [llm-compressor](https://github.com/vllm-project/llm-compressor),
for serving with vLLM.

## What was done

| | |
|---|---|
| Scheme | `FP8_DYNAMIC` — weights static per-channel FP8, activations dynamic per-token FP8 |
| Calibration | none needed; dynamic activation scales are computed at inference time |
| Left in bf16 | `lm_head`, and the token embeddings |
| Weights before | 7.49 GiB |
| Weights after | 4.83 GiB (35% smaller) |

The score is read from the `yes`/`no` logits of the model head, which is left in bf16. Quantising it would put error directly into the number documents are ranked by.

## Quality gate

[`WikipediaRerankingMultilingual`](https://huggingface.co/datasets/mteb/WikipediaRerankingMultilingual) from MTEB — reranking Wikipedia passages in 16 languages, scored by `map_at_1000`.

| language | bf16 | FP8 | Δ |
|---|---|---|---|
| de | 0.9600 | 0.9594 | -0.0006 |
| en | 0.9715 | 0.9726 | +0.0011 |
| it | 0.9710 | 0.9702 | -0.0008 |
| **mean** | **0.9675** | **0.9674** | **-0.0001** |

Languages evaluated: de, en, it. Tolerance: 0.0100 map_at_1000 per language.

**PASSED — no language lost more than 0.0100 map_at_1000.**

## Why

Serving [Qwen/Qwen3-Reranker-4B](https://huggingface.co/Qwen/Qwen3-Reranker-4B) in bf16 leaves little room for KV cache on a small GPU: the weights take
what the cache needs, and the context length has to be cut until it fits. Halving the weights
gives that memory back — the same card serves a longer context without any other change.

## Serving

```bash
vllm serve DCC-BS/Qwen3-Reranker-4B-FP8-Dynamic \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.85
```

The checkpoint is in `compressed-tensors` format, so vLLM detects the quantisation from
`config.json`; no extra flag is required.

FP8 arithmetic is native on Ada and Hopper (compute capability 8.9+). On Ampere it runs through
Marlin: the memory saving still applies, the speed is roughly unchanged.