Instructions to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: llama cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Ollama
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Ollama:
ollama run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Unsloth Studio
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS to start chatting
- Docker Model Runner
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Docker Model Runner:
docker model run hf.co/Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
- Lemonade
How to use Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Clickbook/ClickBook-Gemma-4-E2B-allscripts-IQ4_XS:IQ4_XS
Run and chat with the model
lemonade run user.ClickBook-Gemma-4-E2B-allscripts-IQ4_XS-IQ4_XS
List all available models
lemonade list
- Atomic Chat
| license: apache-2.0 | |
| license_link: https://ai.google.dev/gemma/docs/gemma_4_license | |
| base_model: google/gemma-4-E2B-it | |
| library_name: gguf | |
| pipeline_tag: text-generation | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - on-device | |
| - mobile | |
| - quantized | |
| - multilingual | |
| # ClickBook Gemma 4 E2B β full vocabulary, IQ4_XS | |
| The **maximum script coverage** variant of the ClickBook on-device reading model. | |
| Every script the base model shipped with is retained: no vocabulary pruning. | |
| **1.950 GB.** | |
| > ### Read this before choosing this model | |
| > | |
| > This variant scores **63.6** on our benchmark against **79.2** for the pruned | |
| > [ClickBook-Gemma-4-E2B-multi-IQ4_XS](https://huggingface.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS), | |
| > which is **1.844 GB β smaller as well as better.** | |
| > | |
| > Choose this one only if you need a script the pruned build removes. For every | |
| > language both models support, the pruned build is better in every respect. | |
| ## Why the full vocabulary is worse | |
| Both models are the same weights, quantised identically with the same importance | |
| matrix. The only difference is vocabulary size β 262,144 tokens here against | |
| 231,955 β and it costs 15.6 points. | |
| | build | vocabulary | size | score | failures / 90 | | |
| |---|---:|---:|---:|---:| | |
| | same weights at f16 | 231,955 | 8.676 GB | 80.5 | 12 | | |
| | [pruned, recommended](https://huggingface.co/Clickbook/ClickBook-Gemma-4-E2B-multi-IQ4_XS) | 231,955 | 1.844 GB | **79.2** | 15 | | |
| | narrower prune | 180,850 | 1.666 GB | 79.0 | 15 | | |
| | **this model** | 262,144 | 1.950 GB | **63.6** | 23 | | |
| The regression is worst on the *easiest* third of the benchmark (91.5 β 70.8), | |
| which is the signature of a general degradation rather than a few hard items | |
| going wrong. | |
| A plausible mechanism: at Q2_K precision the embedding table and the output | |
| distribution spend representational capacity on tens of thousands of tokens the | |
| model never needs, at the expense of the tokens it does. **Vocabulary pruning is | |
| not only a size optimisation β it is a quality one.** The measurement is | |
| single-seed and reported as measured; the explanation is a hypothesis. | |
| It shows qualitatively too. Asked to define a word, this build often restates it | |
| ("waiting means to wait for something") or retells the sentence instead of | |
| defining the tapped word β the failure mode the prompts were specifically revised | |
| to eliminate, reappearing here. | |
| ## What this build adds | |
| Fourteen script groups the pruned build removes. Each was probed with a single tap | |
| and **all of them tokenise and return plausible English answers** β no | |
| byte-fallback garbage: | |
| | script | probe result | | |
| |---|---| | |
| | Malayalam | "He waited in front of the shop for a long time." | | |
| | Ethiopic | "He waited for a long time in front of his shop." | | |
| | Georgian | "It was waiting for a long time." | | |
| | Kannada | "He was waiting." | | |
| | Bengali | "Waiting means to stay in one place or time." | | |
| | Telugu, Gujarati, Sinhala | correct but tautological β "waiting means to wait" | | |
| | Lao, Myanmar | correct but very terse β "Wait" | | |
| Also present: Tamil, Thai, Oriya, Gurmukhi, Tibetan, Cherokee, Thaana, Syriac, | |
| Mongolian, NKo and the emoji block. | |
| **One tap per script is not a benchmark.** It establishes that the tokeniser works | |
| and the model is not producing rubbish. It says nothing about quality, and given | |
| the overall 63.6 you should assume quality in these scripts is lower than the | |
| probes suggest. | |
| **Greek remains wrong** here as in the pruned build: asked about `ΟΟΞ¬ΟΡ΢α` used to | |
| mean a park bench, it answered "the bank means a bank". | |
| ## No prompts exist for the new scripts | |
| `prompts.json` in this repo is the same file as the pruned build's and covers | |
| **18 languages only**. The additional scripts here tokenise, but have no MEANING, | |
| EXAMPLE, TRANSLATION or CONTEXT templates. Using them means writing prompts first; | |
| the existing ones are a reasonable model to follow. | |
| ## Usage | |
| Identical to the pruned build. See `MOBILE-INTEGRATION.md`. | |
| ```jsonc | |
| { | |
| "temperature": 0.1, "top_k": 40, "top_p": 0.9, "repeat_penalty": 1.05, | |
| "chat_template_kwargs": { "enable_thinking": false } // REQUIRED | |
| } | |
| ``` | |
| Without `enable_thinking: false` the model spends its whole budget in | |
| `reasoning_content` and returns empty `content` with `finish_reason: "length"` β | |
| indistinguishable from an unsupported language. | |
| ## License and provenance | |
| **Apache License 2.0**, matching the base model, | |
| [`google/gemma-4-E2B-it`](https://huggingface.co/google/gemma-4-E2B-it). Google | |
| also publishes a [Gemma 4 license page](https://ai.google.dev/gemma/docs/gemma_4_license), | |
| linked from the upstream card. | |
| **Modifications**, as Apache 2.0 requires derivative works to state: quantised to | |
| IQ4_XS with Q2_K token embeddings under an importance matrix. **The vocabulary is | |
| unmodified.** No weights were fine-tuned, distilled or retrained. | |
| Gemma is a trademark of Google LLC. This is an independent derivative, not | |
| endorsed by or affiliated with Google. | |