Bonsai-27B-Android-Local / README_EN.md
livadies's picture
Document community value and compare related Android LLM projects / Анализ пользы и похожих проектов
bc577b4 verified
|
Raw
History Blame Contribute Delete
12.6 kB
---
license: other
pipeline_tag: text-generation
language:
- en
- ru
tags:
- android
- gguf
- llama-cpp
- on-device
- bonsai-27b
---
# Bonsai Local for Android
[Русский](README.md) · **English** · [Benchmarks / Замеры](BENCHMARKS.md)
Research Android application that runs
[Prism ML Bonsai-27B GGUF](https://huggingface.co/prism-ml/Bonsai-27B-gguf)
fully on-device through JNI and the PrismML `llama.cpp` fork. Once the GGUF is
copied to the device, chat inference does not use a PC server, cloud API, or
internet connection.
![Bonsai Local with the model loaded](screenshots/02-model-loaded.png)
## What works
- Loads the complete 3,803,452,480-byte `Bonsai-27B-Q1_0.gguf`.
- Performs inference inside the Android application process.
- Streams chat output and exposes the model's thinking trace.
- Auto-discovers GGUF files in private and app-specific external storage.
- Imports compatible GGUF files through Android Storage Access Framework.
- Includes an on-device prompt-processing/token-generation benchmark.
- Builds for physical `arm64-v8a` devices and `x86_64` emulators.
- Includes an adaptive Android launcher icon and raster mipmaps.
- Provides a debug-only Base64 UTF-8 prompt intent for reproducible tests.
The model is deliberately not embedded inside the APK because a 3.8 GB APK is
impractical. A verified copy is published in this repository at
`models/Bonsai-27B-Q1_0.gguf`, with upstream attribution preserved, so the app
and its exact tested model can be retrieved from one place.
## Verified setup
| Component | Value |
|---|---|
| Host | Windows 11, Intel Core Ultra 5 125H |
| Android Studio | 2025.2 |
| AGP / Gradle | 8.13.2 / 8.14.3 |
| Build JDK | JetBrains Runtime 21.0.9 |
| compileSdk / targetSdk / minSdk | 36 / 36 / 30 |
| Android NDK / CMake | 28.2.13676358 / 3.22.1 |
| APK ABIs | `arm64-v8a`, `x86_64` |
| Test AVD | Android 17 preview, x86_64, 8 GB RAM, 16 GB data |
| PrismML fork commit | `62061f91088281e65071cc38c5f69ee95c39f14e` |
## Measured Android results
All measurements below came from the Android emulator and the application JNI
runtime, not from desktop `llama-cli`.
| Check | Result |
|---|---|
| Debug APK build, install, launch | passed |
| Recognized model | `qwen35 27B Q1_0`, 26.9B parameters, 3.53 GiB |
| Backend | CPU, dynamically selected x86_64 variant |
| Prompt processing, pp64 | **2.89 tokens/s** |
| Token generation, tg32 | **1.79 tokens/s** |
| Arithmetic prompt `37 × 19` | correct final answer: **703** |
| Thinking trace | operational |
| Offline inference after import | confirmed |
| Fatal exceptions | none observed |
The Russian/Kotlin instruction-following test also produced a useful negative
result: the model understood the requested language, signature, and validation
constraints, but spent the complete 512-token limit in visible reasoning and
did not reach a clean final code answer. This is recorded rather than presented
as a successful result.
Screenshots:
- [`02-model-loaded.png`](screenshots/02-model-loaded.png) — final installed APK
with the local model loaded;
- [`03-benchmark.png`](screenshots/03-benchmark.png) — on-device benchmark;
- [`04-reasoning.png`](screenshots/04-reasoning.png) — live generation and 703;
- [`05-russian-kotlin.png`](screenshots/05-russian-kotlin.png) — 512-token
thinking-limit result.
Emulator throughput should not be treated as a phone prediction. A physical
device has different SIMD support, thermal limits, memory bandwidth, and power
management.
## Architecture
```mermaid
flowchart TD
UI["Android UI · MainActivity"] --> API["InferenceEngine Kotlin API"]
API --> DISP["Single-thread coroutine dispatcher"]
DISP --> JNI["JNI · libai-chat.so"]
JNI --> COMMON["llama-common · chat template · sampler"]
JNI --> LLAMA["PrismML llama.cpp"]
LLAMA --> GGML["GGML CPU backend loader"]
GGML --> ABI{"Device ABI and CPU features"}
ABI --> ARM["ARM · NEON / DOTPROD / I8MM / SVE / SME"]
ABI --> X86["x86_64 · SSE4 / AVX2 / AVX512 / AMX"]
GGUF["Bonsai-27B-Q1_0.gguf · 3.80 GB"] --> LLAMA
```
Project modules:
- `app/` — UI, model import, chat, benchmark, and Base64 test intent.
- `lib/` — Kotlin API, GGUF metadata reader, JNI, and native build.
- `third_party/llama.cpp/` — PrismML fork, kept next to this project when
building locally and not duplicated in this repository.
- `models/` — the exact GGUF used by the Android tests and its checksum.
- `screenshots/` — evidence captured from the Android emulator.
- `release/` — verified debug APK and checksum.
Model load flow:
1. `MainActivity` waits for the native engine to initialize.
2. It searches private and app-specific external `models` directories.
3. Kotlin passes the absolute GGUF path to `loadModel()`.
4. Native calls are serialized through one IO dispatcher.
5. JNI creates an 8,192-token context, batch size 512, and sampler.
6. GGML selects the best packaged CPU backend for the current ABI/features.
7. The Qwen chat template is applied and token pieces stream as `Flow<String>`.
Critical JNI setup:
```cpp
llama_model_params model_params = llama_model_default_params();
g_model = llama_model_load_from_file(model_path, model_params);
llama_context_params ctx_params = llama_context_default_params();
ctx_params.n_ctx = 8192;
ctx_params.n_batch = 512;
ctx_params.n_threads = n_threads;
g_context = llama_init_from_model(g_model, ctx_params);
```
Native state is serialized because model, context, batch, and sampler are
global native resources:
```kotlin
private val llamaDispatcher = Dispatchers.IO.limitedParallelism(1)
override suspend fun loadModel(pathToModel: String) =
withContext(llamaDispatcher) {
load(pathToModel)
prepare()
}
```
## Runtime capabilities verified
The real full-GGUF load reported:
- 64 transformer blocks and Q1_0 tensors;
- approximately 149.62 MiB recurrent state;
- approximately 523.02 MiB CPU compute buffer;
- automatic Flash Attention;
- fused Gated Delta Net in autoregressive and chunked paths;
- approximately 3,703 graph nodes and one split.
Bonsai-27B is not a Mixture-of-Experts model. It is a dense 27B hybrid-attention
model, so there are no experts or routing network to enumerate. Roughly 75% of
its blocks use linear/recurrent attention and 25% full attention.
This APK is text-only. The optional upstream vision projection
`Bonsai-27B-mmproj-Q8_0.gguf` and DSpark speculative drafter
`Bonsai-27B-dspark-Q4_1.gguf` are not integrated.
## Build
Place the pinned PrismML fork next to this project:
```text
gpt/
├── BonsaiAndroid/
└── third_party/llama.cpp/
```
Install SDK 36, NDK `28.2.13676358`, and CMake `3.22.1`, then open
`BonsaiAndroid` in Android Studio or run:
```powershell
$env:JAVA_HOME = 'C:\Program Files\Android\Android Studio1\jbr'
java -classpath gradle\wrapper\gradle-wrapper.jar `
org.gradle.wrapper.GradleWrapperMain :app:assembleDebug
```
The build output is `app/build/outputs/apk/debug/app-debug.apk`. The tested copy
is published as `release/BonsaiLocal-debug.apk`:
```text
size: 120762843 bytes
SHA256: BDAF2D9EE7EE1BBB2A424242678D75AC35F2B770973E3F4C9E58639CE8F93E5C
```
The native runtime is wired through:
```cmake
set(LLAMA_SRC ${CMAKE_CURRENT_LIST_DIR}/../../../../../third_party/llama.cpp)
add_subdirectory(${LLAMA_SRC} build-llama)
```
## Install the model
Normal path:
1. Copy `Bonsai-27B-Q1_0.gguf` to the Android device.
2. Open Bonsai Local.
3. Tap **Choose GGUF**.
4. Select the file and wait for import/load.
Reproducible emulator path:
```powershell
adb shell mkdir -p `
/sdcard/Android/data/com.prismml.bonsailocal/files/models
adb push .\models\Bonsai-27B-Q1_0.gguf `
/sdcard/Android/data/com.prismml.bonsailocal/files/models/
```
Verified model checksum:
```text
17EF842E47450CAEB8EAA3EBFBBAB5D2F2278B62B79BE107985FB69A2F819AA0
```
## Problems encountered and fixes
1. **No ready Android runtime in the collection.** The upstream Android sample
was adapted and the native runtime is built inside Gradle.
2. **Mainline llama.cpp was insufficient.** Q1_0 and the hybrid architecture
require the pinned PrismML fork.
3. **A separate JDK 17 toolchain was unavailable.** The project builds on JBR
21 while explicitly targeting JVM bytecode 17.
4. **Android logging API level mismatch.** A local priority filter was used and
minSdk was finalized at API 30.
5. **The existing AVD had insufficient RAM/storage.** A dedicated 8 GB RAM,
16 GB data AVD was created without modifying the user's existing Pixel AVD.
6. **16 KB page-size checks.** NDK r28 and AGP 8.13.2 are used; ELF LOAD
alignment is `2**14` and `zipalign -P 16` passes. An unused DataStore bundle
and its `libdatastore_shared_counter.so` were removed. Android 17 preview may
still report experimental RELRO compatibility warnings for dynamically
loaded CPU variants; verify again on stable Android 15/16 before publishing
to production.
7. **A 3.8 GB GGUF is easy to duplicate accidentally.** The test model was
pushed directly into app-specific external storage.
8. **Thinking mode makes short CPU tests take minutes.** Generation remains
asynchronous, but production UI should separate/hide thinking and expose a
configurable token limit.
## Current limitations and next work
- Text-only and CPU-only in the tested AVD.
- No resumable model downloader or in-app SHA-256 verification.
- Chat history is in-memory.
- Application context is currently 8,192 tokens.
- Thinking tags are displayed as normal text.
- 6 GB devices may be killed by Android LMK or run out of memory.
- Vision, Vulkan, compressed KV cache, and DSpark remain research tasks.
Recommended next experiments: physical Snapdragon/Dimensity benchmarks,
first-token latency and thermal tests, Vulkan comparison, collapsible thinking
UI, WorkManager download/resume, vision `mmproj`, speculative decoding, and
per-ABI AAB delivery.
## Community value and related projects
Running an LLM locally on Android is not itself novel. Official `llama.cpp`
already provides an Android Studio sample and runtime CPU-kernel selection;
PocketPal AI and ChatterUI are mature GGUF clients; MLC LLM, MNN, and
ExecuTorch provide alternative Android runtimes and demos.
| Project | Existing capability | Difference in this work |
|---|---|---|
| [llama.cpp Android](https://github.com/ggml-org/llama.cpp/blob/master/docs/android.md) | official JNI/sample and CPU variants | validates this unusual Q1_0 hybrid model and publishes a tested APK |
| [PocketPal AI](https://github.com/a-ghorbani/pocketpal-ai) | mature GGUF client, HF downloads, benchmarks | general product versus a narrow reproducible Bonsai-27B test case |
| [ChatterUI](https://github.com/Vali-98/ChatterUI) | GGUF and API chat through React Native | this project keeps a minimal native Android/JNI path |
| [MLC LLM](https://llm.mlc.ai/docs/deploy/android.html) | GPU-oriented Android SDK and demo | different format/toolchain; this work consumes the original GGUF on CPU |
| [MNN](https://github.com/alibaba/MNN/tree/master/apps/Android/MnnLlmChat) | high-performance multimodal Android client | far broader and faster; this repository is simpler as a Q1_0/GDN regression fixture |
| [ExecuTorch](https://docs.pytorch.org/executorch/stable/llm/run-on-android.html) | AAR and experimental Java LLM API | exports to `.pte`; this project documents the GGUF/llama.cpp route |
A public Hugging Face and GitHub search at publication time did not reveal
another reproducible package for **Bonsai-27B Q1_0 running inside an Android
APK**. This does not prove absolute priority, but the useful combination is:
- the exact GGUF, APK, source, and checksums live together;
- real Android pp/tg measurements, runtime paths, and screenshots are recorded;
- Android preview 16 KB page-size and RELRO issues are documented;
- both a correct reasoning result and a failed Kotlin instruction test are kept;
- the research notes are available in English and Russian.
The current result is best described as an **engineering baseline and regression
artifact**, not a new model architecture or a production competitor to
PocketPal/MNN. Physical ARM64 Snapdragon/Dimensity results, RAM/energy/
first-token measurements, CI builds, and upstream Android fixes would turn it
into a substantially stronger community contribution.
## Licenses
- Bonsai-27B GGUF: Apache-2.0 according to the upstream model card.
- PrismML fork / llama.cpp: see the upstream repository licenses.
- Preserve all applicable upstream `LICENSE` and `NOTICE` files when
redistributing the APK or model.