Instructions to use livadies/Bonsai-27B-Android-Local with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use livadies/Bonsai-27B-Android-Local with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: llama cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: llama cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: ./llama-cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf livadies/Bonsai-27B-Android-Local:Q1_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf livadies/Bonsai-27B-Android-Local:Q1_0
Use Docker
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- LM Studio
- Jan
- vLLM
How to use livadies/Bonsai-27B-Android-Local with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "livadies/Bonsai-27B-Android-Local" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "livadies/Bonsai-27B-Android-Local", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Ollama
How to use livadies/Bonsai-27B-Android-Local with Ollama:
ollama run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Unsloth Studio
How to use livadies/Bonsai-27B-Android-Local with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for livadies/Bonsai-27B-Android-Local to start chatting
- Pi
How to use livadies/Bonsai-27B-Android-Local with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "livadies/Bonsai-27B-Android-Local:Q1_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use livadies/Bonsai-27B-Android-Local with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default livadies/Bonsai-27B-Android-Local:Q1_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use livadies/Bonsai-27B-Android-Local with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf livadies/Bonsai-27B-Android-Local:Q1_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "livadies/Bonsai-27B-Android-Local:Q1_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use livadies/Bonsai-27B-Android-Local with Docker Model Runner:
docker model run hf.co/livadies/Bonsai-27B-Android-Local:Q1_0
- Lemonade
How to use livadies/Bonsai-27B-Android-Local with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull livadies/Bonsai-27B-Android-Local:Q1_0
Run and chat with the model
lemonade run user.Bonsai-27B-Android-Local-Q1_0
List all available models
lemonade list
| license: other | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| - ru | |
| tags: | |
| - android | |
| - gguf | |
| - llama-cpp | |
| - on-device | |
| - bonsai-27b | |
| # Bonsai Local for Android | |
| [Русский](README.md) · **English** · [Benchmarks / Замеры](BENCHMARKS.md) | |
| Research Android application that runs | |
| [Prism ML Bonsai-27B GGUF](https://huggingface.co/prism-ml/Bonsai-27B-gguf) | |
| fully on-device through JNI and the PrismML `llama.cpp` fork. Once the GGUF is | |
| copied to the device, chat inference does not use a PC server, cloud API, or | |
| internet connection. | |
|  | |
| ## What works | |
| - Loads the complete 3,803,452,480-byte `Bonsai-27B-Q1_0.gguf`. | |
| - Performs inference inside the Android application process. | |
| - Streams chat output and exposes the model's thinking trace. | |
| - Auto-discovers GGUF files in private and app-specific external storage. | |
| - Imports compatible GGUF files through Android Storage Access Framework. | |
| - Includes an on-device prompt-processing/token-generation benchmark. | |
| - Builds for physical `arm64-v8a` devices and `x86_64` emulators. | |
| - Includes an adaptive Android launcher icon and raster mipmaps. | |
| - Provides a debug-only Base64 UTF-8 prompt intent for reproducible tests. | |
| The model is deliberately not embedded inside the APK because a 3.8 GB APK is | |
| impractical. A verified copy is published in this repository at | |
| `models/Bonsai-27B-Q1_0.gguf`, with upstream attribution preserved, so the app | |
| and its exact tested model can be retrieved from one place. | |
| ## Verified setup | |
| | Component | Value | | |
| |---|---| | |
| | Host | Windows 11, Intel Core Ultra 5 125H | | |
| | Android Studio | 2025.2 | | |
| | AGP / Gradle | 8.13.2 / 8.14.3 | | |
| | Build JDK | JetBrains Runtime 21.0.9 | | |
| | compileSdk / targetSdk / minSdk | 36 / 36 / 30 | | |
| | Android NDK / CMake | 28.2.13676358 / 3.22.1 | | |
| | APK ABIs | `arm64-v8a`, `x86_64` | | |
| | Test AVD | Android 17 preview, x86_64, 8 GB RAM, 16 GB data | | |
| | PrismML fork commit | `62061f91088281e65071cc38c5f69ee95c39f14e` | | |
| ## Measured Android results | |
| All measurements below came from the Android emulator and the application JNI | |
| runtime, not from desktop `llama-cli`. | |
| | Check | Result | | |
| |---|---| | |
| | Debug APK build, install, launch | passed | | |
| | Recognized model | `qwen35 27B Q1_0`, 26.9B parameters, 3.53 GiB | | |
| | Backend | CPU, dynamically selected x86_64 variant | | |
| | Prompt processing, pp64 | **2.89 tokens/s** | | |
| | Token generation, tg32 | **1.79 tokens/s** | | |
| | Arithmetic prompt `37 × 19` | correct final answer: **703** | | |
| | Thinking trace | operational | | |
| | Offline inference after import | confirmed | | |
| | Fatal exceptions | none observed | | |
| The Russian/Kotlin instruction-following test also produced a useful negative | |
| result: the model understood the requested language, signature, and validation | |
| constraints, but spent the complete 512-token limit in visible reasoning and | |
| did not reach a clean final code answer. This is recorded rather than presented | |
| as a successful result. | |
| Screenshots: | |
| - [`02-model-loaded.png`](screenshots/02-model-loaded.png) — final installed APK | |
| with the local model loaded; | |
| - [`03-benchmark.png`](screenshots/03-benchmark.png) — on-device benchmark; | |
| - [`04-reasoning.png`](screenshots/04-reasoning.png) — live generation and 703; | |
| - [`05-russian-kotlin.png`](screenshots/05-russian-kotlin.png) — 512-token | |
| thinking-limit result. | |
| Emulator throughput should not be treated as a phone prediction. A physical | |
| device has different SIMD support, thermal limits, memory bandwidth, and power | |
| management. | |
| ## Architecture | |
| ```mermaid | |
| flowchart TD | |
| UI["Android UI · MainActivity"] --> API["InferenceEngine Kotlin API"] | |
| API --> DISP["Single-thread coroutine dispatcher"] | |
| DISP --> JNI["JNI · libai-chat.so"] | |
| JNI --> COMMON["llama-common · chat template · sampler"] | |
| JNI --> LLAMA["PrismML llama.cpp"] | |
| LLAMA --> GGML["GGML CPU backend loader"] | |
| GGML --> ABI{"Device ABI and CPU features"} | |
| ABI --> ARM["ARM · NEON / DOTPROD / I8MM / SVE / SME"] | |
| ABI --> X86["x86_64 · SSE4 / AVX2 / AVX512 / AMX"] | |
| GGUF["Bonsai-27B-Q1_0.gguf · 3.80 GB"] --> LLAMA | |
| ``` | |
| Project modules: | |
| - `app/` — UI, model import, chat, benchmark, and Base64 test intent. | |
| - `lib/` — Kotlin API, GGUF metadata reader, JNI, and native build. | |
| - `third_party/llama.cpp/` — PrismML fork, kept next to this project when | |
| building locally and not duplicated in this repository. | |
| - `models/` — the exact GGUF used by the Android tests and its checksum. | |
| - `screenshots/` — evidence captured from the Android emulator. | |
| - `release/` — verified debug APK and checksum. | |
| Model load flow: | |
| 1. `MainActivity` waits for the native engine to initialize. | |
| 2. It searches private and app-specific external `models` directories. | |
| 3. Kotlin passes the absolute GGUF path to `loadModel()`. | |
| 4. Native calls are serialized through one IO dispatcher. | |
| 5. JNI creates an 8,192-token context, batch size 512, and sampler. | |
| 6. GGML selects the best packaged CPU backend for the current ABI/features. | |
| 7. The Qwen chat template is applied and token pieces stream as `Flow<String>`. | |
| Critical JNI setup: | |
| ```cpp | |
| llama_model_params model_params = llama_model_default_params(); | |
| g_model = llama_model_load_from_file(model_path, model_params); | |
| llama_context_params ctx_params = llama_context_default_params(); | |
| ctx_params.n_ctx = 8192; | |
| ctx_params.n_batch = 512; | |
| ctx_params.n_threads = n_threads; | |
| g_context = llama_init_from_model(g_model, ctx_params); | |
| ``` | |
| Native state is serialized because model, context, batch, and sampler are | |
| global native resources: | |
| ```kotlin | |
| private val llamaDispatcher = Dispatchers.IO.limitedParallelism(1) | |
| override suspend fun loadModel(pathToModel: String) = | |
| withContext(llamaDispatcher) { | |
| load(pathToModel) | |
| prepare() | |
| } | |
| ``` | |
| ## Runtime capabilities verified | |
| The real full-GGUF load reported: | |
| - 64 transformer blocks and Q1_0 tensors; | |
| - approximately 149.62 MiB recurrent state; | |
| - approximately 523.02 MiB CPU compute buffer; | |
| - automatic Flash Attention; | |
| - fused Gated Delta Net in autoregressive and chunked paths; | |
| - approximately 3,703 graph nodes and one split. | |
| Bonsai-27B is not a Mixture-of-Experts model. It is a dense 27B hybrid-attention | |
| model, so there are no experts or routing network to enumerate. Roughly 75% of | |
| its blocks use linear/recurrent attention and 25% full attention. | |
| This APK is text-only. The optional upstream vision projection | |
| `Bonsai-27B-mmproj-Q8_0.gguf` and DSpark speculative drafter | |
| `Bonsai-27B-dspark-Q4_1.gguf` are not integrated. | |
| ## Build | |
| Place the pinned PrismML fork next to this project: | |
| ```text | |
| gpt/ | |
| ├── BonsaiAndroid/ | |
| └── third_party/llama.cpp/ | |
| ``` | |
| Install SDK 36, NDK `28.2.13676358`, and CMake `3.22.1`, then open | |
| `BonsaiAndroid` in Android Studio or run: | |
| ```powershell | |
| $env:JAVA_HOME = 'C:\Program Files\Android\Android Studio1\jbr' | |
| java -classpath gradle\wrapper\gradle-wrapper.jar ` | |
| org.gradle.wrapper.GradleWrapperMain :app:assembleDebug | |
| ``` | |
| The build output is `app/build/outputs/apk/debug/app-debug.apk`. The tested copy | |
| is published as `release/BonsaiLocal-debug.apk`: | |
| ```text | |
| size: 120762843 bytes | |
| SHA256: BDAF2D9EE7EE1BBB2A424242678D75AC35F2B770973E3F4C9E58639CE8F93E5C | |
| ``` | |
| The native runtime is wired through: | |
| ```cmake | |
| set(LLAMA_SRC ${CMAKE_CURRENT_LIST_DIR}/../../../../../third_party/llama.cpp) | |
| add_subdirectory(${LLAMA_SRC} build-llama) | |
| ``` | |
| ## Install the model | |
| Normal path: | |
| 1. Copy `Bonsai-27B-Q1_0.gguf` to the Android device. | |
| 2. Open Bonsai Local. | |
| 3. Tap **Choose GGUF**. | |
| 4. Select the file and wait for import/load. | |
| Reproducible emulator path: | |
| ```powershell | |
| adb shell mkdir -p ` | |
| /sdcard/Android/data/com.prismml.bonsailocal/files/models | |
| adb push .\models\Bonsai-27B-Q1_0.gguf ` | |
| /sdcard/Android/data/com.prismml.bonsailocal/files/models/ | |
| ``` | |
| Verified model checksum: | |
| ```text | |
| 17EF842E47450CAEB8EAA3EBFBBAB5D2F2278B62B79BE107985FB69A2F819AA0 | |
| ``` | |
| ## Problems encountered and fixes | |
| 1. **No ready Android runtime in the collection.** The upstream Android sample | |
| was adapted and the native runtime is built inside Gradle. | |
| 2. **Mainline llama.cpp was insufficient.** Q1_0 and the hybrid architecture | |
| require the pinned PrismML fork. | |
| 3. **A separate JDK 17 toolchain was unavailable.** The project builds on JBR | |
| 21 while explicitly targeting JVM bytecode 17. | |
| 4. **Android logging API level mismatch.** A local priority filter was used and | |
| minSdk was finalized at API 30. | |
| 5. **The existing AVD had insufficient RAM/storage.** A dedicated 8 GB RAM, | |
| 16 GB data AVD was created without modifying the user's existing Pixel AVD. | |
| 6. **16 KB page-size checks.** NDK r28 and AGP 8.13.2 are used; ELF LOAD | |
| alignment is `2**14` and `zipalign -P 16` passes. An unused DataStore bundle | |
| and its `libdatastore_shared_counter.so` were removed. Android 17 preview may | |
| still report experimental RELRO compatibility warnings for dynamically | |
| loaded CPU variants; verify again on stable Android 15/16 before publishing | |
| to production. | |
| 7. **A 3.8 GB GGUF is easy to duplicate accidentally.** The test model was | |
| pushed directly into app-specific external storage. | |
| 8. **Thinking mode makes short CPU tests take minutes.** Generation remains | |
| asynchronous, but production UI should separate/hide thinking and expose a | |
| configurable token limit. | |
| ## Current limitations and next work | |
| - Text-only and CPU-only in the tested AVD. | |
| - No resumable model downloader or in-app SHA-256 verification. | |
| - Chat history is in-memory. | |
| - Application context is currently 8,192 tokens. | |
| - Thinking tags are displayed as normal text. | |
| - 6 GB devices may be killed by Android LMK or run out of memory. | |
| - Vision, Vulkan, compressed KV cache, and DSpark remain research tasks. | |
| Recommended next experiments: physical Snapdragon/Dimensity benchmarks, | |
| first-token latency and thermal tests, Vulkan comparison, collapsible thinking | |
| UI, WorkManager download/resume, vision `mmproj`, speculative decoding, and | |
| per-ABI AAB delivery. | |
| ## Community value and related projects | |
| Running an LLM locally on Android is not itself novel. Official `llama.cpp` | |
| already provides an Android Studio sample and runtime CPU-kernel selection; | |
| PocketPal AI and ChatterUI are mature GGUF clients; MLC LLM, MNN, and | |
| ExecuTorch provide alternative Android runtimes and demos. | |
| | Project | Existing capability | Difference in this work | | |
| |---|---|---| | |
| | [llama.cpp Android](https://github.com/ggml-org/llama.cpp/blob/master/docs/android.md) | official JNI/sample and CPU variants | validates this unusual Q1_0 hybrid model and publishes a tested APK | | |
| | [PocketPal AI](https://github.com/a-ghorbani/pocketpal-ai) | mature GGUF client, HF downloads, benchmarks | general product versus a narrow reproducible Bonsai-27B test case | | |
| | [ChatterUI](https://github.com/Vali-98/ChatterUI) | GGUF and API chat through React Native | this project keeps a minimal native Android/JNI path | | |
| | [MLC LLM](https://llm.mlc.ai/docs/deploy/android.html) | GPU-oriented Android SDK and demo | different format/toolchain; this work consumes the original GGUF on CPU | | |
| | [MNN](https://github.com/alibaba/MNN/tree/master/apps/Android/MnnLlmChat) | high-performance multimodal Android client | far broader and faster; this repository is simpler as a Q1_0/GDN regression fixture | | |
| | [ExecuTorch](https://docs.pytorch.org/executorch/stable/llm/run-on-android.html) | AAR and experimental Java LLM API | exports to `.pte`; this project documents the GGUF/llama.cpp route | | |
| A public Hugging Face and GitHub search at publication time did not reveal | |
| another reproducible package for **Bonsai-27B Q1_0 running inside an Android | |
| APK**. This does not prove absolute priority, but the useful combination is: | |
| - the exact GGUF, APK, source, and checksums live together; | |
| - real Android pp/tg measurements, runtime paths, and screenshots are recorded; | |
| - Android preview 16 KB page-size and RELRO issues are documented; | |
| - both a correct reasoning result and a failed Kotlin instruction test are kept; | |
| - the research notes are available in English and Russian. | |
| The current result is best described as an **engineering baseline and regression | |
| artifact**, not a new model architecture or a production competitor to | |
| PocketPal/MNN. Physical ARM64 Snapdragon/Dimensity results, RAM/energy/ | |
| first-token measurements, CI builds, and upstream Android fixes would turn it | |
| into a substantially stronger community contribution. | |
| ## Licenses | |
| - Bonsai-27B GGUF: Apache-2.0 according to the upstream model card. | |
| - PrismML fork / llama.cpp: see the upstream repository licenses. | |
| - Preserve all applicable upstream `LICENSE` and `NOTICE` files when | |
| redistributing the APK or model. | |