Instructions to use livadies/Bonsai-27B-Android-Local with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use livadies/Bonsai-27B-Android-Local with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: llama cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: llama cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: ./llama-cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Use Docker
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- LM Studio
- Jan
- vLLM
How to use livadies/Bonsai-27B-Android-Local with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "livadies/Bonsai-27B-Android-Local" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "livadies/Bonsai-27B-Android-Local", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Ollama
How to use livadies/Bonsai-27B-Android-Local with Ollama:
ollama run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Unsloth Studio
How to use livadies/Bonsai-27B-Android-Local with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
- Pi
How to use livadies/Bonsai-27B-Android-Local with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "livadies/Bonsai-27B-Android-Local:Q1_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use livadies/Bonsai-27B-Android-Local with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default livadies/Bonsai-27B-Android-Local:Q1_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use livadies/Bonsai-27B-Android-Local with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "livadies/Bonsai-27B-Android-Local:Q1_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use livadies/Bonsai-27B-Android-Local with Docker Model Runner:
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Lemonade
How to use livadies/Bonsai-27B-Android-Local with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull livadies/Bonsai-27B-Android-Local:Q1_0
Run and chat with the model
lemonade run user.Bonsai-27B-Android-Local-Q1_0
List all available models
lemonade list
File size: 12,635 Bytes
9863dbd a32665e 9863dbd a32665e 9863dbd a32665e 9863dbd bc577b4 9863dbd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 | ---
license: other
pipeline_tag: text-generation
language:
- en
- ru
tags:
- android
- gguf
- llama-cpp
- on-device
- bonsai-27b
---
# Bonsai Local for Android
[Русский](README.md) · **English** · [Benchmarks / Замеры](BENCHMARKS.md)
Research Android application that runs
[Prism ML Bonsai-27B GGUF](https://huggingface.co/prism-ml/Bonsai-27B-gguf)
fully on-device through JNI and the PrismML `llama.cpp` fork. Once the GGUF is
copied to the device, chat inference does not use a PC server, cloud API, or
internet connection.

## What works
- Loads the complete 3,803,452,480-byte `Bonsai-27B-Q1_0.gguf`.
- Performs inference inside the Android application process.
- Streams chat output and exposes the model's thinking trace.
- Auto-discovers GGUF files in private and app-specific external storage.
- Imports compatible GGUF files through Android Storage Access Framework.
- Includes an on-device prompt-processing/token-generation benchmark.
- Builds for physical `arm64-v8a` devices and `x86_64` emulators.
- Includes an adaptive Android launcher icon and raster mipmaps.
- Provides a debug-only Base64 UTF-8 prompt intent for reproducible tests.
The model is deliberately not embedded inside the APK because a 3.8 GB APK is
impractical. A verified copy is published in this repository at
`models/Bonsai-27B-Q1_0.gguf`, with upstream attribution preserved, so the app
and its exact tested model can be retrieved from one place.
## Verified setup
| Component | Value |
|---|---|
| Host | Windows 11, Intel Core Ultra 5 125H |
| Android Studio | 2025.2 |
| AGP / Gradle | 8.13.2 / 8.14.3 |
| Build JDK | JetBrains Runtime 21.0.9 |
| compileSdk / targetSdk / minSdk | 36 / 36 / 30 |
| Android NDK / CMake | 28.2.13676358 / 3.22.1 |
| APK ABIs | `arm64-v8a`, `x86_64` |
| Test AVD | Android 17 preview, x86_64, 8 GB RAM, 16 GB data |
| PrismML fork commit | `62061f91088281e65071cc38c5f69ee95c39f14e` |
## Measured Android results
All measurements below came from the Android emulator and the application JNI
runtime, not from desktop `llama-cli`.
| Check | Result |
|---|---|
| Debug APK build, install, launch | passed |
| Recognized model | `qwen35 27B Q1_0`, 26.9B parameters, 3.53 GiB |
| Backend | CPU, dynamically selected x86_64 variant |
| Prompt processing, pp64 | **2.89 tokens/s** |
| Token generation, tg32 | **1.79 tokens/s** |
| Arithmetic prompt `37 × 19` | correct final answer: **703** |
| Thinking trace | operational |
| Offline inference after import | confirmed |
| Fatal exceptions | none observed |
The Russian/Kotlin instruction-following test also produced a useful negative
result: the model understood the requested language, signature, and validation
constraints, but spent the complete 512-token limit in visible reasoning and
did not reach a clean final code answer. This is recorded rather than presented
as a successful result.
Screenshots:
- [`02-model-loaded.png`](screenshots/02-model-loaded.png) — final installed APK
with the local model loaded;
- [`03-benchmark.png`](screenshots/03-benchmark.png) — on-device benchmark;
- [`04-reasoning.png`](screenshots/04-reasoning.png) — live generation and 703;
- [`05-russian-kotlin.png`](screenshots/05-russian-kotlin.png) — 512-token
thinking-limit result.
Emulator throughput should not be treated as a phone prediction. A physical
device has different SIMD support, thermal limits, memory bandwidth, and power
management.
## Architecture
```mermaid
flowchart TD
UI["Android UI · MainActivity"] --> API["InferenceEngine Kotlin API"]
API --> DISP["Single-thread coroutine dispatcher"]
DISP --> JNI["JNI · libai-chat.so"]
JNI --> COMMON["llama-common · chat template · sampler"]
JNI --> LLAMA["PrismML llama.cpp"]
LLAMA --> GGML["GGML CPU backend loader"]
GGML --> ABI{"Device ABI and CPU features"}
ABI --> ARM["ARM · NEON / DOTPROD / I8MM / SVE / SME"]
ABI --> X86["x86_64 · SSE4 / AVX2 / AVX512 / AMX"]
GGUF["Bonsai-27B-Q1_0.gguf · 3.80 GB"] --> LLAMA
```
Project modules:
- `app/` — UI, model import, chat, benchmark, and Base64 test intent.
- `lib/` — Kotlin API, GGUF metadata reader, JNI, and native build.
- `third_party/llama.cpp/` — PrismML fork, kept next to this project when
building locally and not duplicated in this repository.
- `models/` — the exact GGUF used by the Android tests and its checksum.
- `screenshots/` — evidence captured from the Android emulator.
- `release/` — verified debug APK and checksum.
Model load flow:
1. `MainActivity` waits for the native engine to initialize.
2. It searches private and app-specific external `models` directories.
3. Kotlin passes the absolute GGUF path to `loadModel()`.
4. Native calls are serialized through one IO dispatcher.
5. JNI creates an 8,192-token context, batch size 512, and sampler.
6. GGML selects the best packaged CPU backend for the current ABI/features.
7. The Qwen chat template is applied and token pieces stream as `Flow<String>`.
Critical JNI setup:
```cpp
llama_model_params model_params = llama_model_default_params();
g_model = llama_model_load_from_file(model_path, model_params);
llama_context_params ctx_params = llama_context_default_params();
ctx_params.n_ctx = 8192;
ctx_params.n_batch = 512;
ctx_params.n_threads = n_threads;
g_context = llama_init_from_model(g_model, ctx_params);
```
Native state is serialized because model, context, batch, and sampler are
global native resources:
```kotlin
private val llamaDispatcher = Dispatchers.IO.limitedParallelism(1)
override suspend fun loadModel(pathToModel: String) =
withContext(llamaDispatcher) {
load(pathToModel)
prepare()
}
```
## Runtime capabilities verified
The real full-GGUF load reported:
- 64 transformer blocks and Q1_0 tensors;
- approximately 149.62 MiB recurrent state;
- approximately 523.02 MiB CPU compute buffer;
- automatic Flash Attention;
- fused Gated Delta Net in autoregressive and chunked paths;
- approximately 3,703 graph nodes and one split.
Bonsai-27B is not a Mixture-of-Experts model. It is a dense 27B hybrid-attention
model, so there are no experts or routing network to enumerate. Roughly 75% of
its blocks use linear/recurrent attention and 25% full attention.
This APK is text-only. The optional upstream vision projection
`Bonsai-27B-mmproj-Q8_0.gguf` and DSpark speculative drafter
`Bonsai-27B-dspark-Q4_1.gguf` are not integrated.
## Build
Place the pinned PrismML fork next to this project:
```text
gpt/
├── BonsaiAndroid/
└── third_party/llama.cpp/
```
Install SDK 36, NDK `28.2.13676358`, and CMake `3.22.1`, then open
`BonsaiAndroid` in Android Studio or run:
```powershell
$env:JAVA_HOME = 'C:\Program Files\Android\Android Studio1\jbr'
java -classpath gradle\wrapper\gradle-wrapper.jar `
org.gradle.wrapper.GradleWrapperMain :app:assembleDebug
```
The build output is `app/build/outputs/apk/debug/app-debug.apk`. The tested copy
is published as `release/BonsaiLocal-debug.apk`:
```text
size: 120762843 bytes
SHA256: BDAF2D9EE7EE1BBB2A424242678D75AC35F2B770973E3F4C9E58639CE8F93E5C
```
The native runtime is wired through:
```cmake
set(LLAMA_SRC ${CMAKE_CURRENT_LIST_DIR}/../../../../../third_party/llama.cpp)
add_subdirectory(${LLAMA_SRC} build-llama)
```
## Install the model
Normal path:
1. Copy `Bonsai-27B-Q1_0.gguf` to the Android device.
2. Open Bonsai Local.
3. Tap **Choose GGUF**.
4. Select the file and wait for import/load.
Reproducible emulator path:
```powershell
adb shell mkdir -p `
/sdcard/Android/data/com.prismml.bonsailocal/files/models
adb push .\models\Bonsai-27B-Q1_0.gguf `
/sdcard/Android/data/com.prismml.bonsailocal/files/models/
```
Verified model checksum:
```text
17EF842E47450CAEB8EAA3EBFBBAB5D2F2278B62B79BE107985FB69A2F819AA0
```
## Problems encountered and fixes
1. **No ready Android runtime in the collection.** The upstream Android sample
was adapted and the native runtime is built inside Gradle.
2. **Mainline llama.cpp was insufficient.** Q1_0 and the hybrid architecture
require the pinned PrismML fork.
3. **A separate JDK 17 toolchain was unavailable.** The project builds on JBR
21 while explicitly targeting JVM bytecode 17.
4. **Android logging API level mismatch.** A local priority filter was used and
minSdk was finalized at API 30.
5. **The existing AVD had insufficient RAM/storage.** A dedicated 8 GB RAM,
16 GB data AVD was created without modifying the user's existing Pixel AVD.
6. **16 KB page-size checks.** NDK r28 and AGP 8.13.2 are used; ELF LOAD
alignment is `2**14` and `zipalign -P 16` passes. An unused DataStore bundle
and its `libdatastore_shared_counter.so` were removed. Android 17 preview may
still report experimental RELRO compatibility warnings for dynamically
loaded CPU variants; verify again on stable Android 15/16 before publishing
to production.
7. **A 3.8 GB GGUF is easy to duplicate accidentally.** The test model was
pushed directly into app-specific external storage.
8. **Thinking mode makes short CPU tests take minutes.** Generation remains
asynchronous, but production UI should separate/hide thinking and expose a
configurable token limit.
## Current limitations and next work
- Text-only and CPU-only in the tested AVD.
- No resumable model downloader or in-app SHA-256 verification.
- Chat history is in-memory.
- Application context is currently 8,192 tokens.
- Thinking tags are displayed as normal text.
- 6 GB devices may be killed by Android LMK or run out of memory.
- Vision, Vulkan, compressed KV cache, and DSpark remain research tasks.
Recommended next experiments: physical Snapdragon/Dimensity benchmarks,
first-token latency and thermal tests, Vulkan comparison, collapsible thinking
UI, WorkManager download/resume, vision `mmproj`, speculative decoding, and
per-ABI AAB delivery.
## Community value and related projects
Running an LLM locally on Android is not itself novel. Official `llama.cpp`
already provides an Android Studio sample and runtime CPU-kernel selection;
PocketPal AI and ChatterUI are mature GGUF clients; MLC LLM, MNN, and
ExecuTorch provide alternative Android runtimes and demos.
| Project | Existing capability | Difference in this work |
|---|---|---|
| [llama.cpp Android](https://github.com/ggml-org/llama.cpp/blob/master/docs/android.md) | official JNI/sample and CPU variants | validates this unusual Q1_0 hybrid model and publishes a tested APK |
| [PocketPal AI](https://github.com/a-ghorbani/pocketpal-ai) | mature GGUF client, HF downloads, benchmarks | general product versus a narrow reproducible Bonsai-27B test case |
| [ChatterUI](https://github.com/Vali-98/ChatterUI) | GGUF and API chat through React Native | this project keeps a minimal native Android/JNI path |
| [MLC LLM](https://llm.mlc.ai/docs/deploy/android.html) | GPU-oriented Android SDK and demo | different format/toolchain; this work consumes the original GGUF on CPU |
| [MNN](https://github.com/alibaba/MNN/tree/master/apps/Android/MnnLlmChat) | high-performance multimodal Android client | far broader and faster; this repository is simpler as a Q1_0/GDN regression fixture |
| [ExecuTorch](https://docs.pytorch.org/executorch/stable/llm/run-on-android.html) | AAR and experimental Java LLM API | exports to `.pte`; this project documents the GGUF/llama.cpp route |
A public Hugging Face and GitHub search at publication time did not reveal
another reproducible package for **Bonsai-27B Q1_0 running inside an Android
APK**. This does not prove absolute priority, but the useful combination is:
- the exact GGUF, APK, source, and checksums live together;
- real Android pp/tg measurements, runtime paths, and screenshots are recorded;
- Android preview 16 KB page-size and RELRO issues are documented;
- both a correct reasoning result and a failed Kotlin instruction test are kept;
- the research notes are available in English and Russian.
The current result is best described as an **engineering baseline and regression
artifact**, not a new model architecture or a production competitor to
PocketPal/MNN. Physical ARM64 Snapdragon/Dimensity results, RAM/energy/
first-token measurements, CI builds, and upstream Android fixes would turn it
into a substantially stronger community contribution.
## Licenses
- Bonsai-27B GGUF: Apache-2.0 according to the upstream model card.
- PrismML fork / llama.cpp: see the upstream repository licenses.
- Preserve all applicable upstream `LICENSE` and `NOTICE` files when
redistributing the APK or model.
|