gemma-4-31B ยท FP8 (dynamic)

8-bit quantization of google/gemma-4-31B (commit 5bbc2fb), produced for efficient vLLM serving. Quantized by axteam.

Base model google/gemma-4-31B (commit 5bbc2fb)
Method FP8 (dynamic)
Format compressed-tensors FP8_DYNAMIC
Weights / activations W8A8 โ€” FP8 E4M3 weights (per-output-channel scale), dynamic per-token FP8 activations
Architecture Gemma 4 ยท Dense
Size on disk ~31.0 GB
Calibration None (data-free static weight scales)
License Gemma 4 / Apache-2.0

Quantization details

  • Scheme โ€” W8A8 โ€” FP8 E4M3 weights (per-output-channel scale), dynamic per-token FP8 activations, stored as compressed-tensors FP8_DYNAMIC.
  • Kept in BF16 โ€” vision tower, vision projection, token embeddings (tied with lm_head), MoE router, and all norms. These are left unquantized to preserve quality (a small fraction of total parameters).
  • Tokenizer / chat_template / processor_config / generation_config are byte-identical to the original, including Google's July-2026 corrected chat_template. config.json is the original plus a quantization_config block.

Evaluation

Quantization fidelity was checked against the BF16 original using perplexity and top-1 next-token agreement on an internal mixed English/Korean set. In our runs the 8-bit variants (AWQ INT8 / FP8) stay closest to the original, while the 4-bit variants (AWQ INT4 / NVFP4) trade a little quality for roughly half the footprint.

Usage (vLLM)

vllm serve axteam/gemma-4-31B-FP8-dynamic
from vllm import LLM, SamplingParams

llm = LLM(model="axteam/gemma-4-31B-FP8-dynamic")
params = SamplingParams(temperature=1.0, top_p=0.95, top_k=64, max_tokens=1024)
out = llm.chat(
    [{"role": "user", "content": "Write a short joke about saving RAM."}],
    params,
)
print(out[0].outputs[0].text)

Recommended sampling (per Google): temperature=1.0, top_p=0.95, top_k=64. Thinking mode: add the <|think|> token at the start of the system prompt to enable step-by-step reasoning.

Reducing KV-cache memory (optional, serve-time)

The KV cache is unquantized (BF16) here โ€” the standard for quantized checkpoints (RedHatAI, NVIDIA, etc. ship the same way). KV-cache precision is a runtime choice, independent of the weight quantization above. For long context or large batches you can roughly halve KV-cache memory with FP8:

vllm serve axteam/gemma-4-31B-FP8-dynamic --kv-cache-dtype fp8

--kv-cache-dtype accepts auto (default, BF16), fp8 / fp8_e4m3, fp8_e5m2. For a GGUF build the llama.cpp equivalent is --cache-type-k q8_0 --cache-type-v q8_0.

License

This is a quantized derivative of Google's Gemma 4. Use is governed by the Gemma 4 license. All original weights, tokenizer, and chat template belong to Google DeepMind.


ํ•œ๊ตญ์–ด ์š”์•ฝ

google/gemma-4-31B ๋ฅผ FP8 (dynamic) ๋กœ ์–‘์žํ™”ํ•œ ์„œ๋น™์šฉ(8-bit) ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค. ์ž‘์„ฑ: axteam.

  • ๋ฐฉ์‹ โ€” FP8 E4M3. ๊ฐ€์ค‘์น˜๋Š” ์ถœ๋ ฅ ์ฑ„๋„๋งˆ๋‹ค ์Šค์ผ€์ผ ํ•˜๋‚˜(์ •์ ยท๋ณด์ • ๋ถˆํ•„์š”), ํ™œ์„ฑํ™”๋Š” ์ถ”๋ก  ์‹œ ํ† ํฐ๋งˆ๋‹ค ๋™์  FP8. ํ˜•์‹์€ compressed-tensors FP8_DYNAMIC.
  • BF16 ์œ ์ง€ โ€” ๋น„์ „ ํƒ€์›Œยท๋น„์ „ ํˆฌ์˜ยทํ† ํฐ ์ž„๋ฒ ๋”ฉ(lm_head์™€ ๋ฌถ์ž„)ยทMoE ๋ผ์šฐํ„ฐยท๋…ธ๋ฆ„ (์ „์ฒด์˜ ๊ทนํžˆ ์ผ๋ถ€).
  • ํ† ํฌ๋‚˜์ด์ €ยทchat_templateยทprocessorยทgeneration_config ๋Š” ์›๋ณธ๊ณผ ๋ฐ”์ดํŠธ๊นŒ์ง€ ๋™์ผ (๊ตฌ๊ธ€์ด 2026-07์— ๊ณ ์นœ chat_template ๊ทธ๋Œ€๋กœ). config.json ์€ ์›๋ณธ + quantization_config.
  • ํ’ˆ์งˆ โ€” BF16 ์›๋ณธ ๋Œ€๋น„ perplexityยทtop-1 ์ผ์น˜๋„๋กœ ๊ฒ€์ฆ(์˜/ํ•œ ํ˜ผํ•ฉ์…‹). 8๋น„ํŠธ(AWQ INT8ยทFP8)๊ฐ€ ์›๋ณธ์— ๊ฐ€์žฅ ๊ฐ€๊น๊ณ , 4๋น„ํŠธ(AWQ INT4ยทNVFP4)๋Š” ์šฉ๋Ÿ‰์„ ์ ˆ๋ฐ˜์œผ๋กœ ์ค„์ด๋Š” ๋Œ€์‹  ํ’ˆ์งˆ์„ ์กฐ๊ธˆ ์–‘๋ณด.
  • ์‚ฌ์šฉ โ€” vllm serve axteam/gemma-4-31B-FP8-dynamic. ๊ถŒ์žฅ ์ƒ˜ํ”Œ๋ง temperature=1.0, top_p=0.95, top_k=64.
  • KV ์บ์‹œ ์ ˆ๊ฐ(์„ ํƒ, ์„œ๋น™ ์‹œ) โ€” ๊ธฐ๋ณธ์€ BF16 KV(์–‘์žํ™” ์ฒดํฌํฌ์ธํŠธ ํ‘œ์ค€, RedHatยทNVIDIA ๋™์ผ). ๊ฐ€์ค‘์น˜ ์–‘์žํ™”์™€ ๋ฌด๊ด€ํ•œ ๋Ÿฐํƒ€์ž„ ์„ ํƒ. ๊ธด ๋ฌธ๋งฅยท๋Œ€๋ฐฐ์น˜์—์„  vllm serve โ€ฆ --kv-cache-dtype fp8 ๋กœ KV ๋ฉ”๋ชจ๋ฆฌ ์ ˆ๋ฐ˜. (GGUF๋ผ๋ฉด llama.cpp --cache-type-k q8_0 --cache-type-v q8_0)
  • ๋ผ์ด์„ ์Šค โ€” ๊ตฌ๊ธ€ Gemma 4 ํŒŒ์ƒ๋ฌผ. Gemma 4 ๋ผ์ด์„ ์Šค ์ ์šฉ. ์›๋ณธ ๊ฐ€์ค‘์น˜ยทํ† ํฌ๋‚˜์ด์ €ยทchat_template ๊ถŒ๋ฆฌ๋Š” Google DeepMind.
Downloads last month
18
Safetensors
Model size
31B params
Tensor type
BF16
ยท
F8_E4M3
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for axteam/gemma-4-31B-FP8-dynamic

Quantized
(46)
this model

Collection including axteam/gemma-4-31B-FP8-dynamic