gemma-2-2b — FP8_DYNAMIC (W8A8-e4m3)

Weight-FP8 checkpoint of google/gemma-2-2b-it, produced for the TR171 deployment-time safety-tax benchmark.

Provenance

Field Value
Base model google/gemma-2-2b-it
Base revision 299a8560bedf22ed1c72a8a11e7dce4a7f9f51f8 (verified — recorded in the frozen matrix)
Recipe FP8_DYNAMIC (W8A8-e4m3), llmcompressor
Quantization method compressed-tensors
Calibration data none — FP8_DYNAMIC is data-free
Build date 2026-07-02
Shard size 3.24 GB
Quantize wall time 177.5 s
Integrity record per-file sha256 from Hub LFS metadata; shard_bytes verified

Reproducing

Producer: research/tr171/expansion/fp8_support_probe.py; environment: research/tr171/expansion/Dockerfile.fp8. The recipe takes no calibration corpus, so there is no dataset or seed to reproduce — only the base checkpoint and the toolchain version.

Known reproducibility gap: llmcompressor was unpinned at build time, so the exact version used on 2026-07-02 is unrecorded. The Dockerfile now pins it. A rebuild may therefore not be bit-identical to this artifact.

Integrity, stated honestly: the 2026-07-02 build recorded no sha256 of its own, and the local build directory is now empty, so no aggregate directory digest exists for it. What is verifiable instead: the per-file sha256 below is read from this repo's Git-LFS metadata, and the mirror was checked against the build record — summing the file sizes in this repo, excluding the generated README.md, NOTICE and .gitattributes, reproduces the matrix's shard_bytes of 3,240,242,075 exactly. So these hashes describe the same bytes the probe measured, and you can verify a download against them directly:

File sha256
model.safetensors 05a93e8dcf0bc822adcf7985510c5635946dc72f49b3f6a356641515d82522d1
tokenizer.json 487cee8724215dcd2dde8888539e8b1bf844ceb5dbbe27f7845abda69eeb060f

LFS-tracked files only; the small JSON/text files are git blobs and carry no sha256. fp8_support_probe.py now records a real digest at build time, so future shards will not need this reconstruction.

License and notices

Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms. Use is subject to the Gemma Prohibited Use Policy at ai.google.dev/gemma/prohibited_use_policy. These terms travel with this derivative and must be passed on to any downstream recipient.

This FP8 derivative inherits the upstream terms of google/gemma-2-2b-it. Consult the base model's licence before redistributing.

Downloads last month
10
Safetensors
Model size
3B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Crusadersk/gemma-2-2b-it-FP8-Dynamic-TR171

Quantized
(193)
this model

Collection including Crusadersk/gemma-2-2b-it-FP8-Dynamic-TR171