Instructions to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Ollama
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Ollama:
ollama run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Unsloth Studio
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
- Docker Model Runner
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Docker Model Runner:
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Lemonade
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Run and chat with the model
lemonade run user.ClickBook-Gemma-4-E2B-allscripts-IQ4_XS-IQ4_XS
List all available models
lemonade list
- Atomic Chat
File size: 4,948 Bytes
55d6529 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 | ---
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: google/gemma-4-E2B-it
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- on-device
- mobile
- quantized
- multilingual
---
# ClickBook Gemma 4 E2B — full vocabulary, IQ4_XS
The **maximum script coverage** variant of the ClickBook on-device reading model.
Every script the base model shipped with is retained: no vocabulary pruning.
**1.950 GB.**
> ### Read this before choosing this model
>
> This variant scores **63.6** on our benchmark against **79.2** for the pruned
> [ClickBook-Gemma-4-E2B-multi-IQ4_XS](https://huggingface.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS),
> which is **1.844 GB — smaller as well as better.**
>
> Choose this one only if you need a script the pruned build removes. For every
> language both models support, the pruned build is better in every respect.
## Why the full vocabulary is worse
Both models are the same weights, quantised identically with the same importance
matrix. The only difference is vocabulary size — 262,144 tokens here against
231,955 — and it costs 15.6 points.
| build | vocabulary | size | score | failures / 90 |
|---|---:|---:|---:|---:|
| same weights at f16 | 231,955 | 8.676 GB | 80.5 | 12 |
| [pruned, recommended](https://huggingface.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS) | 231,955 | 1.844 GB | **79.2** | 15 |
| narrower prune | 180,850 | 1.666 GB | 79.0 | 15 |
| **this model** | 262,144 | 1.950 GB | **63.6** | 23 |
The regression is worst on the *easiest* third of the benchmark (91.5 → 70.8),
which is the signature of a general degradation rather than a few hard items
going wrong.
A plausible mechanism: at Q2_K precision the embedding table and the output
distribution spend representational capacity on tens of thousands of tokens the
model never needs, at the expense of the tokens it does. **Vocabulary pruning is
not only a size optimisation — it is a quality one.** The measurement is
single-seed and reported as measured; the explanation is a hypothesis.
It shows qualitatively too. Asked to define a word, this build often restates it
("waiting means to wait for something") or retells the sentence instead of
defining the tapped word — the failure mode the prompts were specifically revised
to eliminate, reappearing here.
## What this build adds
Fourteen script groups the pruned build removes. Each was probed with a single tap
and **all of them tokenise and return plausible English answers** — no
byte-fallback garbage:
| script | probe result |
|---|---|
| Malayalam | "He waited in front of the shop for a long time." |
| Ethiopic | "He waited for a long time in front of his shop." |
| Georgian | "It was waiting for a long time." |
| Kannada | "He was waiting." |
| Bengali | "Waiting means to stay in one place or time." |
| Telugu, Gujarati, Sinhala | correct but tautological — "waiting means to wait" |
| Lao, Myanmar | correct but very terse — "Wait" |
Also present: Tamil, Thai, Oriya, Gurmukhi, Tibetan, Cherokee, Thaana, Syriac,
Mongolian, NKo and the emoji block.
**One tap per script is not a benchmark.** It establishes that the tokeniser works
and the model is not producing rubbish. It says nothing about quality, and given
the overall 63.6 you should assume quality in these scripts is lower than the
probes suggest.
**Greek remains wrong** here as in the pruned build: asked about `τράπεζα` used to
mean a park bench, it answered "the bank means a bank".
## No prompts exist for the new scripts
`prompts.json` in this repo is the same file as the pruned build's and covers
**18 languages only**. The additional scripts here tokenise, but have no MEANING,
EXAMPLE, TRANSLATION or CONTEXT templates. Using them means writing prompts first;
the existing ones are a reasonable model to follow.
## Usage
Identical to the pruned build. See `MOBILE-INTEGRATION.md`.
```jsonc
{
"temperature": 0.1, "top_k": 40, "top_p": 0.9, "repeat_penalty": 1.05,
"chat_template_kwargs": { "enable_thinking": false } // REQUIRED
}
```
Without `enable_thinking: false` the model spends its whole budget in
`reasoning_content` and returns empty `content` with `finish_reason: "length"` —
indistinguishable from an unsupported language.
## License and provenance
**Apache License 2.0**, matching the base model,
[`google/gemma-4-E2B-it`](https://huggingface.co/google/gemma-4-E2B-it). Google
also publishes a [Gemma 4 license page](https://ai.google.dev/gemma/docs/gemma_4_license),
linked from the upstream card.
**Modifications**, as Apache 2.0 requires derivative works to state: quantised to
IQ4_XS with Q2_K token embeddings under an importance matrix. **The vocabulary is
unmodified.** No weights were fine-tuned, distilled or retrained.
Gemma is a trademark of Google LLC. This is an independent derivative, not
endorsed by or affiliated with Google.
|