--- license: gemma base_model: google/gemma-4-e4b-it tags: - quantized - gguf - candle - q8_0 --- # gemma-4-e4b-it — Q8_0 weight cache + auxiliary tensors Quantized deployment artifacts for [google/gemma-4-e4b-it](https://huggingface.co/google/gemma-4-e4b-it). - `qcache-q8_0.v1.gguf` — GGML **Q8_0** tensors for the model's linear projections, stored in a GGUF container keyed by the original checkpoint tensor paths (no metadata KVs). Produced by in-memory quantization of the bf16 checkpoint with a [candle](https://github.com/huggingface/candle)-based loader. - `aux-tensors.safetensors` — everything the cache does not carry: embeddings, norms and other non-projection tensors, in their original dtypes. - Config and tokenizer files are copied unmodified from the base repository. **Format note**: this is *not* a llama.cpp-compatible GGUF — tensors keep their original checkpoint names and only linear projections are quantized. Load it with a runtime that pairs the cache with the auxiliary safetensors. Weights are redistributed under the same terms as the base model.