Gemma 4 E4B Nightcap — MLX 8-bit

MLX quantization of andyoneal/Gemma-4-E4B-Nightcap, a roleplay / creative-writing task-arithmetic merge of Gemma 4 E4B. See the main repo for the full card — what's in it, how it behaves, and how the recipe was chosen.

Text-only (the conversion keeps the language stack); chat template included.

variant bits/weight size notes
this repo 8.501 ~7.4 GB ~8 GB peak memory, ~50 tok/s on Apple silicon
uv run --with mlx-lm python -m mlx_lm generate \
  --model andyoneal/Gemma-4-E4B-Nightcap-MLX-8bit \
  --system-prompt "You are..." --prompt "..." \
  --chat-template-config '{"enable_thinking": false}' --temp 1.0

Samplers: temp 1.0, min-p 0.02, no repetition penalties (they do more harm than good on this family). Temp 0.8 / top-p 0.95 also works fine.

Downloads last month
59
Safetensors
Model size
7B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for andyoneal/Gemma-4-E4B-Nightcap-MLX-8bit

Quantized
(2)
this model