ThinkLess-2B-FP8 / README.md
Shaik1903's picture
Card: AWQ note moved below How to use; size 2.5 GB
5af8995 verified
|
Raw History Blame Contribute Delete
2.56 kB
metadata
license: apache-2.0
base_model: Shaik1903/ThinkLess-2B
base_model_relation: quantized
language:
  - en
pipeline_tag: text-generation
library_name: transformers
tags:
  - reasoning
  - efficient-reasoning
  - math
  - fp8
  - compressed-tensors
  - qwen3.5

ThinkLess-2B-FP8

FP8 version of ThinkLess-2B: 8-bit floating-point weights and activations, 2.5 GB (bf16: 4.3 GB), with near-identical accuracy. Made with llm-compressor (FP8_DYNAMIC: per-channel FP8 weights, dynamic per-token FP8 activations, no calibration data). The output head, vision tower and MTP heads stay in 16-bit.

Accuracy (81,920-token budget, thinking on)

Benchmark ThinkLess-2B (bf16) ThinkLess-2B-FP8 Mean tokens: bf16 β†’ FP8
GSM8K 90.1 88.6 3,341 β†’ 3,512
MATH-500 88.8 88.2 12,412 β†’ 12,680
GPQA-Diamond 52.8 51.5 16,370 β†’ 17,270

The differences are within the 95% confidence intervals, and answers stay just as short (cut-offs ≀ 1%).

Serving (vLLM 0.30, one H100, max 8,192 output tokens)

Median latency per request on vLLM

Configuration Concurrency 1: tokens/s Concurrency 1: median latency Concurrency 16: requests/s MTP acceptance
Qwen3.5-2B (base) 400 20.0 s 0.66 –
ThinkLess-2B (bf16) 396 10.3 s 0.83 –
ThinkLess-2B-FP8 440 9.6 s 0.88 –
ThinkLess-2B-FP8 + MTP 557 6.9 s 0.99 54%

How to use

vllm serve Shaik1903/ThinkLess-2B-FP8 --speculative-config '{"method":"mtp","num_speculative_tokens":2}'

FP8 compute needs a GPU with FP8 support (NVIDIA Hopper or Ada, e.g. H100, L4, RTX 40-series); vLLM loads the compressed-tensors format directly. Use Qwen3.5's thinking-mode sampling (temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5).

Why FP8 rather than 4-bit

A 4-bit AWQ version of ThinkLess-2B was also evaluated: it lost 7–15 points (MATH-500 88.8 β†’ 74.1), made answers longer and was slower than bf16 on an H100. Small reasoning models are sensitive to low-bit weights over long reasoning chains; FP8 keeps the accuracy.

Accuracy and answer length after compression: bf16 vs FP8 vs AWQ 4-bit

Training details, evaluation protocol and limitations: ThinkLess-2B.