🗜️ DeepGemma-2B-Reasoning — GGUF (Q4_K_M)

This is the quantized GGUF version of Zhantas/DeepGemma-2B-Reasoning.

For the full model, training details and benchmarks — see the original repo.

📦 File

File Size Description
gemma4_e2b-q4_k_m.gguf ~4.7 GB Q4_K_M quantization

⚡ Performance (RTX 4090, llama.cpp, ngl=999)

Metric Speed
Prompt processing ~400 tok/s
Generation ~239–262 tok/s
Context 4096 tokens

💻 Usage

./llama-cli -m gemma4_e2b-q4_k_m.gguf \
  -p "<start_of_turn>user\nYour question here<end_of_turn>\n<start_of_turn>model\n" \
  -n 512 -ngl 999 -c 4096

⚠️ Limitations

Prone to "overthinking" on simple tasks. Best suited for logic puzzles, coding, and mathematics.

Downloads last month
6
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Zhantas/DeepGemma-2B-Reasoning-GGUF

Quantized
(1)
this model