Instructions to use axteam/gemma-4-31B-FP8-dynamic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use axteam/gemma-4-31B-FP8-dynamic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="axteam/gemma-4-31B-FP8-dynamic")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("axteam/gemma-4-31B-FP8-dynamic") model = AutoModelForMultimodalLM.from_pretrained("axteam/gemma-4-31B-FP8-dynamic", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use axteam/gemma-4-31B-FP8-dynamic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "axteam/gemma-4-31B-FP8-dynamic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axteam/gemma-4-31B-FP8-dynamic", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/axteam/gemma-4-31B-FP8-dynamic
- SGLang
How to use axteam/gemma-4-31B-FP8-dynamic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "axteam/gemma-4-31B-FP8-dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axteam/gemma-4-31B-FP8-dynamic", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "axteam/gemma-4-31B-FP8-dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "axteam/gemma-4-31B-FP8-dynamic", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use axteam/gemma-4-31B-FP8-dynamic with Docker Model Runner:
docker model run hf.co/axteam/gemma-4-31B-FP8-dynamic
gemma-4-31B ยท FP8 (dynamic)
8-bit quantization of google/gemma-4-31B (commit 5bbc2fb), produced for efficient vLLM serving. Quantized by axteam.
| Base model | google/gemma-4-31B (commit 5bbc2fb) |
| Method | FP8 (dynamic) |
| Format | compressed-tensors FP8_DYNAMIC |
| Weights / activations | W8A8 โ FP8 E4M3 weights (per-output-channel scale), dynamic per-token FP8 activations |
| Architecture | Gemma 4 ยท Dense |
| Size on disk | ~31.0 GB |
| Calibration | None (data-free static weight scales) |
| License | Gemma 4 / Apache-2.0 |
Quantization details
- Scheme โ W8A8 โ FP8 E4M3 weights (per-output-channel scale), dynamic per-token FP8 activations, stored as compressed-tensors
FP8_DYNAMIC. - Kept in BF16 โ vision tower, vision projection, token embeddings (tied with lm_head), MoE router, and all norms. These are left unquantized to preserve quality (a small fraction of total parameters).
- Tokenizer /
chat_template/processor_config/generation_configare byte-identical to the original, including Google's July-2026 correctedchat_template.config.jsonis the original plus aquantization_configblock.
Evaluation
Quantization fidelity was checked against the BF16 original using perplexity and top-1 next-token agreement on an internal mixed English/Korean set. In our runs the 8-bit variants (AWQ INT8 / FP8) stay closest to the original, while the 4-bit variants (AWQ INT4 / NVFP4) trade a little quality for roughly half the footprint.
Usage (vLLM)
vllm serve axteam/gemma-4-31B-FP8-dynamic
from vllm import LLM, SamplingParams
llm = LLM(model="axteam/gemma-4-31B-FP8-dynamic")
params = SamplingParams(temperature=1.0, top_p=0.95, top_k=64, max_tokens=1024)
out = llm.chat(
[{"role": "user", "content": "Write a short joke about saving RAM."}],
params,
)
print(out[0].outputs[0].text)
Recommended sampling (per Google):
temperature=1.0,top_p=0.95,top_k=64. Thinking mode: add the<|think|>token at the start of the system prompt to enable step-by-step reasoning.
Reducing KV-cache memory (optional, serve-time)
The KV cache is unquantized (BF16) here โ the standard for quantized checkpoints (RedHatAI, NVIDIA, etc. ship the same way). KV-cache precision is a runtime choice, independent of the weight quantization above. For long context or large batches you can roughly halve KV-cache memory with FP8:
vllm serve axteam/gemma-4-31B-FP8-dynamic --kv-cache-dtype fp8
--kv-cache-dtypeacceptsauto(default, BF16),fp8/fp8_e4m3,fp8_e5m2. For a GGUF build the llama.cpp equivalent is--cache-type-k q8_0 --cache-type-v q8_0.
License
This is a quantized derivative of Google's Gemma 4. Use is governed by the Gemma 4 license. All original weights, tokenizer, and chat template belong to Google DeepMind.
ํ๊ตญ์ด ์์ฝ
google/gemma-4-31B ๋ฅผ FP8 (dynamic) ๋ก ์์ํํ ์๋น์ฉ(8-bit) ๋ชจ๋ธ์
๋๋ค. ์์ฑ: axteam.
- ๋ฐฉ์ โ FP8 E4M3. ๊ฐ์ค์น๋ ์ถ๋ ฅ ์ฑ๋๋ง๋ค ์ค์ผ์ผ ํ๋(์ ์ ยท๋ณด์ ๋ถํ์), ํ์ฑํ๋ ์ถ๋ก ์ ํ ํฐ๋ง๋ค ๋์ FP8. ํ์์ compressed-tensors
FP8_DYNAMIC. - BF16 ์ ์ง โ ๋น์ ํ์ยท๋น์ ํฌ์ยทํ ํฐ ์๋ฒ ๋ฉ(lm_head์ ๋ฌถ์)ยทMoE ๋ผ์ฐํฐยท๋ ธ๋ฆ (์ ์ฒด์ ๊ทนํ ์ผ๋ถ).
- ํ ํฌ๋์ด์ ยทchat_templateยทprocessorยทgeneration_config ๋ ์๋ณธ๊ณผ ๋ฐ์ดํธ๊น์ง ๋์ผ (๊ตฌ๊ธ์ด 2026-07์ ๊ณ ์น chat_template ๊ทธ๋๋ก).
config.json์ ์๋ณธ +quantization_config. - ํ์ง โ BF16 ์๋ณธ ๋๋น perplexityยทtop-1 ์ผ์น๋๋ก ๊ฒ์ฆ(์/ํ ํผํฉ์ ). 8๋นํธ(AWQ INT8ยทFP8)๊ฐ ์๋ณธ์ ๊ฐ์ฅ ๊ฐ๊น๊ณ , 4๋นํธ(AWQ INT4ยทNVFP4)๋ ์ฉ๋์ ์ ๋ฐ์ผ๋ก ์ค์ด๋ ๋์ ํ์ง์ ์กฐ๊ธ ์๋ณด.
- ์ฌ์ฉ โ
vllm serve axteam/gemma-4-31B-FP8-dynamic. ๊ถ์ฅ ์ํ๋งtemperature=1.0, top_p=0.95, top_k=64. - KV ์บ์ ์ ๊ฐ(์ ํ, ์๋น ์) โ ๊ธฐ๋ณธ์ BF16 KV(์์ํ ์ฒดํฌํฌ์ธํธ ํ์ค, RedHatยทNVIDIA ๋์ผ). ๊ฐ์ค์น ์์ํ์ ๋ฌด๊ดํ ๋ฐํ์ ์ ํ. ๊ธด ๋ฌธ๋งฅยท๋๋ฐฐ์น์์
vllm serve โฆ --kv-cache-dtype fp8๋ก KV ๋ฉ๋ชจ๋ฆฌ ์ ๋ฐ. (GGUF๋ผ๋ฉด llama.cpp--cache-type-k q8_0 --cache-type-v q8_0) - ๋ผ์ด์ ์ค โ ๊ตฌ๊ธ Gemma 4 ํ์๋ฌผ. Gemma 4 ๋ผ์ด์ ์ค ์ ์ฉ. ์๋ณธ ๊ฐ์ค์นยทํ ํฌ๋์ด์ ยทchat_template ๊ถ๋ฆฌ๋ Google DeepMind.
- Downloads last month
- 18
Model tree for axteam/gemma-4-31B-FP8-dynamic
Base model
google/gemma-4-31B