Gemma 4 E2B GGUF (4-bit)

Model Description

This repository contains the Gemma 4 E2B model quantized to 4-bit GGUF format using Unsloth and llama.cpp. Gemma 4 E2B is an extremely efficient 2B parameter model (with approximately 2B effective parameters) designed for high performance on edge devices and low-latency applications.

Quantization Details

  • Quantization Format: GGUF (q4_k_m)
  • Quantization Method: llama.cpp / Unsloth
  • Precision: 4-bit
  • Efficiency: Optimized for local inference with Ollama, LM Studio, and llama.cpp.

Use with Ollama

You can run this model directly using Ollama:

ollama run hf.co/DuoNeural/Gemma-4-E2B-GGUF

Use with LM Studio

  1. Open LM Studio.
  2. Search for DuoNeural/Gemma-4-E2B-GGUF.
  3. Download the Q4_K_M version and load it.

Architecture

Gemma 4 E2B is part of Google's latest lightweight model family, featuring state-of-the-art attention and architecture improvements that allow it to punch far above its weight class in coding and general reasoning.

Limitations

  • Performance may be limited for extremely long-form creative writing or highly complex multi-step logical puzzles compared to larger Gemma 4 variants.
  • Not recommended for tasks requiring high-precision floating-point arithmetic.

DuoNeural

DuoNeural is an open AI research lab — human + AI in collaboration.

🤗 HuggingFace huggingface.co/DuoNeural
🐙 GitHub github.com/DuoNeural
🐦 X / Twitter @DuoNeural
📧 Email duoneural@proton.me
📬 Newsletter duoneural.beehiiv.com
☕ Support buymeacoffee.com/duoneural
🌐 Site duoneural.com

Research Team

  • Jesse — Vision, hardware, direction
  • Archon — AI lab partner, post-training, abliteration, experiments
  • Aura — Research AI, literature synthesis, novel proposals

Raw updates from the lab: model drops, training results, findings. Subscribe at duoneural.beehiiv.com.

DuoNeural Research Publications

Open access, CC BY 4.0. Authored by Archon, Jesse Caldwell, Aura — DuoNeural.

Downloads last month
235
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DuoNeural/Gemma-4-E2B-GGUF

Quantized
(308)
this model