Instructions to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
- Ollama
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Ollama:
ollama run hf.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
- Unsloth Studio
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS to start chatting
- Pi
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Docker Model Runner:
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
- Lemonade
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Run and chat with the model
lemonade run user.ClickBook-Gemma-4-E2B-multi-IQ4_XS-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS:IQ4_XS
Run Hermes
hermes
- Atomic Chat
ClickBook on-device model — integration guide
For the iOS and Android teams. Every number here is measured on this repo's benchmark or its ship gate; where something is untested it says so explicitly.
Start here: ship models/wide_iq4xs_e2.gguf and out/prompts.json. Everything
else in this document explains those two files.
1. The two shipping assets
models/wide_iq4xs_e2.gguf — 1.844 GB
| base | Gemma 4 E2B instruction-tuned |
| quantisation | IQ4_XS body, Q2_K token embeddings, imatrix-guided |
| vocabulary | 231,955 tokens (of 262,144; the rest are pruned scripts) |
| runtime | llama.cpp b10326 or later |
| benchmark | 79.2 on the 90-item set, 15 failures, no empty or truncated panels |
That 79.2 is measured on this exact file and is level with the 1.666 GB eleven-language build (79.0 at the same seed, 78.4 averaged over two). Restoring seven scripts costs nothing on the languages already served — a earlier intermediate build read 1.4 marks lower, which a second seed showed to be noise.
out/prompts.json — 32 KB
Every prompt, per language, with {{word}}, {{sentence}} and {{passage}}
placeholders, plus the required request settings and per-language token caps.
Substitute into these strings; do not retype or paraphrase them. The wording is the result of seven measured revisions and small changes have moved benchmark scores by double digits. Regenerate the file after any prompt change:
BENCH_PROMPTS=v7 npx tsx scripts/export-prompts.ts > out/prompts.json
Other builds, same recipe, if the size budget changes
| file | size | scripts |
|---|---|---|
iq4xs_e2.gguf |
1.666 GB | Latin, Cyrillic, Arabic |
cjkkrhi_iq4xs_e2.gguf |
1.815 GB | + CJK, Hangul, Devanagari |
wide_iq4xs_e2.gguf |
1.844 GB | + Tamil, Thai, Greek, Hebrew |
allscripts_iq4xs_e2.gguf |
1.950 GB | every script — over a 1.9 GB cap |
Vocabulary costs 3.49 KB per token: Gemma 4 keeps per-layer embeddings, so a
token carries 10,496 parameters. Script coverage is therefore expensive, and the
shipped subset is deliberate. scripts/build-script-variant.sh builds any other
combination.
2. Language support
out/prompts.json carries book, reader and evidence for every language.
Honour those flags — they are ship gates, not preferences.
| language | book | reader | evidence |
|---|---|---|---|
| English, German, Arabic | ✅ | ✅ | benchmark, 90 items × 2 seeds |
| French, Portuguese, Spanish, Italian, Russian, Dutch, Polish, Turkish | ✅ | — | benchmark, 183 items each on an older build: 69–86 |
| Chinese, Japanese, Korean, Hindi, Thai, Hebrew | ✅ | ✅ | ship gate: 4/4 tabs pass |
| Tamil | ✅ | ❌ | ship gate: CONTEXT fails |
| Greek | ❌ | ❌ | ship gate: output corrupted |
Reader-language support for French through Turkish needs TRANSLATION and CONTEXT templates, which do not exist yet. Those languages can be read from, not answered in.
Greek — do not enable
Output mixes scripts and invents words:
"η σκέψη του μυаλу" Cyrillic а and у inside a Greek word
"που άμεσως führte" a German verb
"για να β Olympic" nonsense
Only 1,511 Greek tokens were retained, and that is not enough. Leaving them in the file is harmless; just never offer Greek as a book or reader language.
Tamil — book language only
MEANING, EXAMPLE and TRANSLATION pass. CONTEXT does not: translating a passage
into Tamil produced அவர்店の앞 அதிக நேரம்… — Han, Kana and Hangul spliced into
Tamil. A Tamil speaker can read a Tamil book with English, German or Arabic
answers; do not offer Tamil as the answer language.
The six smoke-tested languages
Each was probed with taps where the obvious reading is wrong. All six read them correctly:
| language | tap | got |
|---|---|---|
| Chinese | 不是个东西 | a scornful insult, not "a thing" |
| Japanese | 手が空いた | became free, not "hand is empty" |
| Korean | 발이 넓다 | well connected, not "wide feet" |
| Hindi | हाथ धोना | to lose something, not "wash hands" |
| Thai | ใจดี | kind-hearted |
| Hebrew | יד חמה | a knack for it, not "warm hand" |
Two taps per language is not a benchmark. English, German and Arabic have 90 items across two seeds; these have a handful of sentences. Treat them as beta and expect to find failures the gate does not cover.
3. Request configuration
{
"n_ctx": 2048, // 555 suffices for one tap and is faster
"n_threads": 6,
"n_gpu_layers": 99, // 0 on CPU-only devices
"temperature": 0.1,
"top_k": 40,
"top_p": 0.9,
"repeat_penalty": 1.05,
"seed": 42,
"chat_template_kwargs": { "enable_thinking": false } // REQUIRED
}
enable_thinking: false is not optional
Without it the model reasons before answering, spends the whole budget in
reasoning_content, and returns empty content with finish_reason: "length".
In the UI that is indistinguishable from an unsupported language or a broken
model. It has cost this project three separate debugging sessions, including one
where nine languages appeared to fail and none of them had.
If you ever enable thinking deliberately, raise every cap below to 1000 first.
Token caps
Defaults, and the per-language values in prompts.json under caps.perLanguage:
| tab | default | ko / th | hi | ta | he |
|---|---|---|---|---|---|
| MEANING | 111 | 136 | 144 | 160 | 192 |
| EXAMPLE | 222 | 272 | 288 | 320 | 384 |
| TRANSLATION | 111 | 136 | 144 | 160 | 192 |
| CONTEXT | 444 | 544 | 576 | 640 | 768 |
Chinese and Japanese use the defaults — measured, they cost the same tokens as
English for equivalent content (scripts/token-cost.ts). Hebrew is the most
expensive script at 1.67×.
Verified headroom at these caps: the longest MEANING answer across all seven new languages was 25 of 112 tokens, and no tab in any language came within 10% of its cap. Nothing truncates.
Performance
12.1 s per popup on the device configuration (CPU, ctx 555, all four tabs); roughly 1.7 s per tab with GPU offload.
4. How a tap becomes four answers
| tab | prompt & answer language | asks for |
|---|---|---|
| MEANING | book | define the tapped word in the sense this sentence gives it |
| EXAMPLE | book | three new short sentences using that same sense |
| TRANSLATION | reader | what the word means, answered in the reader's language |
| CONTEXT | reader | translate the whole passage |
Two rules that carry measured weight:
{{sentence}}must come from the tap's character offsets, not from splitting the passage on punctuation. When the same word appears twice in a passage, the wrong sentence produces a confidently wrong answer.{{passage}}is the whole block, or roughly 520 characters centred on the tap when the block is longer.
The idiom clause inside each prompt — "if the word is part of an idiom, explain the whole phrase instead" — is what makes idiom taps work. Do not trim it for brevity.
5. Verifying a build before release
npx tsx scripts/newlang-verify.ts models/wide_iq4xs_e2.gguf
Exits non-zero if any language fails. It checks, per tab: non-empty output, no foreign-script contamination, inside the token cap with headroom, and that EXAMPLE returns three unnumbered lines. Current state: 27 of 28 checks pass, the exception being Tamil CONTEXT, which is why Tamil is book-only.
To read the answers rather than assert on them:
npx tsx scripts/newlang-probe.ts models/wide_iq4xs_e2.gguf zh,ja,ko
For the eleven benchmarked languages:
SENSE10_MODEL=models/wide_iq4xs_e2.gguf bash scripts/full90.sh myarm
6. What is not in this build
There is no on-device dictionary. A purpose-built sense resource measured
+12.7 marks (79.0 → 91.7 on the 90-item benchmark) but needs a generation run that
has not been funded; the version buildable from free public lexicons measured
+0.7 and is not worth its 82.5 MB. Format and integration points are in
docs/sense10.md. wide_iq4xs_e2.gguf leaves 56 MB under a 1.9 GB cap if one
ships later.
Known open issues, both reproducible with the ship gate:
- Tamil CONTEXT script contamination (§2).
- EXAMPLE occasionally numbers its sentences despite the prompt forbidding it.
Strip a leading
1./-/•client-side.
7. Packaging and delivery
The model is delivered by the stores themselves — Play Asset Delivery on Android,
Background Assets on iOS. Both host and serve it free, both resume interrupted
downloads, and both tie the asset to an app version, which matters here because
prompts.json describes gates and caps for this model file. A plain CDN URL
lets those drift apart; the store mechanisms do not.
Integrity
file ClickBook-Gemma-4-E2B-multi-IQ4_XS.gguf
bytes 1,844,170,944
sha256 0ad915db36b04976d4a217e264d5a46c2d08fb77632875795d7ac080b8d5f6fe
Verify after assembly, before first load. A truncated or mis-ordered file will often still load in llama.cpp and produce fluent nonsense — the failure does not announce itself.
Android — Play Asset Delivery, two packs
The per-pack limit is 1.5 GB and this file is 1.844 GB, so it must be split. Note the limit applies to compressed download size, which does not help: a GGUF is quantised weights and compresses by only 3% (ratio 0.970, measured), so compressed size is effectively file size.
Every variant of this model exceeds 1.5 GB, including the 1.666 GB eleven-language build, so splitting is unavoidable and model size is therefore not a packaging consideration when choosing which build to ship.
Use fast-follow delivery: the app opens immediately and the model arrives in the background behind a progress UI you control. Two packs of 922.1 MB, split at the exact byte midpoint and concatenated in order:
split -b 922085472 ClickBook-Gemma-4-E2B-multi-IQ4_XS.gguf model.part
# model.partaa bytes 0..922,085,471
# sha256 d98cccc0847c9f2200d2eeba1f0d928d70f945528d08c52abfcd1231c7a003b1
# model.partab bytes 922,085,472..1,844,170,943
# sha256 d09e50a639d5f3d28eae17d9708217f9fedfb240063ec3dade5bcf6b46b43393
Two packs sit far inside the 30 GB cumulative fast-follow allowance, so there is no second limit to design around.
iOS — Background Assets, one pack
On-Demand Resources is deprecated as of iOS 27. Build on Background Assets with Apple-hosted asset packs, which allow 200 GB across up to 200 packs.
The per-pack limit on iOS 18+ is 8 GB, so the model ships as a single 1.844 GB pack with no splitting and no reassembly. On iOS 17 and earlier the limit is 512 MB per tag, which would mean four parts — if the minimum deployment target is below iOS 18, raise it or accept the extra client complexity.
Consequence: the two platforms need different packaging
Android reassembles two parts; iOS receives one file. Do not design one chunking scheme for both. Keep the split and concatenation logic on the Android side only, and let iOS treat the model as an ordinary single asset.
| limit per pack | parts for 1.844 GB | |
|---|---|---|
| Android, Play Asset Delivery | 1.5 GB | 2 |
| iOS 18+, Background Assets | 8 GB | 1 |
| iOS ≤ 17, On-Demand Resources | 512 MB | 4 (and deprecated) |
8. Release checklist
Everything below has been run against the assets in this repo. Re-run after any change to the model, the prompts or the caps.
| # | check | command | current |
|---|---|---|---|
| 1 | model is the expected size | ls -l models/wide_iq4xs_e2.gguf |
1.844 GB ✅ |
| 2 | no regression on the benchmarked languages | SENSE10_MODEL=models/wide_iq4xs_e2.gguf bash scripts/full90.sh rel |
79.2 vs 79.0 ✅ |
| 3 | new languages pass the ship gate | npx tsx scripts/newlang-verify.ts models/wide_iq4xs_e2.gguf |
27/28 ✅ (Tamil CONTEXT known) |
| 4 | prompts export is current and at v7 | BENCH_PROMPTS=v7 npx tsx scripts/export-prompts.ts > out/prompts.json |
32 KB, v7 ✅ |
| 5 | client sends enable_thinking: false |
inspect one request | verify per platform |
| 6 | client honours book / reader flags |
Greek hidden, Tamil not offered as answer language | verify per platform |
| 7 | client derives {{sentence}} from tap offsets |
not from punctuation splitting | verify per platform |
| 8 | assembled model matches the checksum | sha256sum after concatenation |
verify per platform |
Rows 5–8 cannot be checked from this repo and are the ones most likely to be got wrong. Two of them fail silently: row 5 looks like a model defect, and row 8 — a mis-assembled model — loads without complaint and generates fluent nonsense.
Reproducing every number in this document
bash scripts/full90.sh <arm> # 90-item benchmark
npx tsx scripts/newlang-verify.ts <model.gguf> # ship gate, exits non-zero
npx tsx scripts/newlang-probe.ts <model.gguf> zh,ja,ko # read the answers
npx tsx scripts/token-cost.ts <model.gguf> # token cost per script
bash scripts/build-script-variant.sh <name> <Scripts,Comma> # a different script set