POCKET-26B-CPU / README.md
SeaWolf-AI's picture
POCKET family cross-links + Image/Studio/Zimage
c508118 verified
|
Raw
History Blame Contribute Delete
2.91 kB
metadata
title: POCKET-26B vs Bonsai · CPU (Gemma4-based)
emoji: ⚔️
colorFrom: green
colorTo: yellow
sdk: docker
app_port: 7860
pinned: true
license: apache-2.0
models:
  - google/gemma-4-26B-A4B-it
  - FINAL-Bench/POCKET-26B-GGUF
  - prism-ml/Bonsai-27B-gguf
short_description: BONSAI vs POCKET  26B Gemma4 MoE out-runs a 27B on CPU

BONSAI vs POCKET-26B · live A/B on a CPU

A live demo of POCKET-26B (Q4_K_M, 17 GB) answering on an upgraded CPU Space — no GPU — with a second tab that races it head-to-head against Bonsai-27B (the most-downloaded 1-bit on-device model) on the same box, same stock llama.cpp. Bonsai answers first, then POCKET — watch the tok/s.

POCKET-26B is the Gemma4 sibling of the POCKET family: built from Google's Gemma4-26B-A4B (Apache-2.0), a sparse Mixture-of-Experts model (25.2B total, ~4B active/token). It is unpruned and keeps GPQA-Diamond 67% — on par with the full base — via Korean-imatrix calibration and mixed-precision quantization. Runs on stock llama.cpp: on a 12 GB phone, and on a PC with no graphics card.

Built with FastAPI + prebuilt llama.cpp. Apache-2.0 · a VIDRAFT model family. This Space runs on CPU hardware only.


🧩 The POCKET Family — On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

📚 Full POCKET collection