> The .litertlm models on this card already use the QAT that is discussed in the blog post. The most popular file, `gemma-4-E2B-it.litertlm`, uses a mixture of int2, int4 and int8 to keep it small, fast and efficient. https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/discussions/30#6a2c2e1ac539f50158631020 # Voice Chatbot (Android) — interruptible, on-device voice chat pipeline A Kotlin/Compose Android app implementing a fully local, **interruptible** voice chatbot: ``` mic ─► Silero VAD (sherpa-onnx) │ endpointing + barge-in ▼ Whisper tiny.en Q4 (whisper.cpp) ──► transcript ▼ LLM (streaming): ├─ Local: LiteRT-LM (`.litertlm`, e.g. Qwen3-0.6B) └─ Remote: OpenAI-compatible `/v1/chat/completions` (SSE) ▼ SentenceSplitter ──► per-sentence VITS TTS (sherpa-onnx + Piper voice) ▼ AudioTrack (playback; barge-in stops it instantly) ``` **Barge-in:** speak while the assistant is talking → playback and generation stop immediately and your new utterance becomes the next turn. ## Components | Piece | Runtime | Model | Size | |---|---|---|---| | STT | whisper.cpp (JNI, built via NDK/CMake) | `ggml-tiny.en-q4_0.bin` (Q4) | ~45 MB | | VAD | sherpa-onnx AAR (official prebuilt `sherpa-onnx-1.13.4.aar`) | `silero_vad.onnx` | ~2 MB | | TTS | sherpa-onnx VITS (Piper) | `vits-piper-en_US-lessac-medium` (+ `espeak-ng-data`) | ~62 MB | | LLM (local) | LiteRT-LM `com.google.ai.edge.litertlm:litertlm-android:0.15.0` | `litert-community/Qwen3-0.6B` → `Qwen3-0.6B.litertlm` (dynamic INT8) | 586 MB | | LLM (remote) | OkHttp + kotlinx.serialization (OpenAI-compatible chat completions) | any remote model, e.g. `qwen2.5-0.5b-instruct` | n/a | All models are Apache-2.0 or MIT except the Piper voice/GPL espeak-ng data (bundled in the sherpa-onnx tts-model tarball) — fine for local testing. ## Prerequisites - macOS/Linux with Android SDK (platform 35, build-tools 36, NDK `27.2.12479018`, CMake 3.22.1), JDK 17, and a connected Android device (arm64, 8 GB RAM recommended). - `adb` on PATH. ## Build Toolchain: **AGP 8.13.2 · Gradle 8.13 (wrapper) · Kotlin 2.3.21 · JDK 17**. whisper.cpp is vendored under `third_party/` (committed), so the project builds out-of-the-box once the SDK/NDK are present. ```bash # 0. JDK 17 (Gradle 8.13 does not support newer JDKs) # macOS/Homebrew: brew install openjdk@17 export JAVA_HOME=/opt/homebrew/opt/openjdk@17/libexec/openjdk.jdk/Contents/Home # 0b. Point Gradle at the Android SDK (or set ANDROID_HOME) echo "sdk.dir=$HOME/Library/Android/sdk" > local.properties # 1. (first time) install NDK/cmake if missing "$ANDROID_HOME"/cmdline-tools/latest/bin/sdkmanager \ "platforms;android-35" "ndk;27.2.12479018" "cmake;3.22.1" # 2. Build (Gradle downloads the sherpa-onnx AAR into app/libs on first run) ./gradlew :app:assembleDebug # 3. Install adb install -r app/build/outputs/apk/debug/app-debug.apk ``` Unit tests: `./gradlew :app:testDebugUnitTest` > Only if `third_party/whisper.cpp` is ever missing (e.g. a fresh export that > dropped it): `tools/fetch-deps.sh` re-clones it. ## Models **Models are bundled in the APK** (`assets/models/`) and copied to the app's private internal storage on first launch — no adb push needed. Storage used on the device: `~240 MB` under `/data/data/com.example.voicechatbot/files/models/`. > The APK is ~500 MB because of the bundled models; devices need **~1.5 GB free** > to install it. **Preinstalled:** SmolLM2 135M (LLM) + Whisper tiny.en Q4 + Silero VAD + Piper TTS. **Download in-app:** bigger LLMs (SmolLM2 360M, Qwen3 0.6B, Gemma 4 E2B/E4B/12B) and Whisper models are selected in Settings via **preset chips that auto-download the file if it isn't already saved** (progress shown on the chip). | id | Model | Size | Preinstalled | Thinking | License | |---|---|---|---|---|---| | `smollm2-135m` | SmolLM2 135M | 136 MB | ✅ | no | Apache-2.0 | | `smollm2-360m` | SmolLM2 360M | 356 MB | download | no | Apache-2.0 | | `qwen3-0.6b` | Qwen3 0.6B int8 | 586 MB | download | yes (stripped) | Apache-2.0 | | `gemma-4-e2b` | Gemma 4 E2B (mobile 2/4/8-bit) | 2583 MB | download | yes (stripped) | Apache-2.0 | | `gemma-4-e4b` | Gemma 4 E4B (mobile 2/4/8-bit) | 3654 MB | download | yes (stripped) | Apache-2.0 | | `gemma-4-12b` | Gemma 4 12B | 6548 MB | download | yes (stripped) | Apache-2.0 | To rebuild with bundled models, the files must be present in `./models/`: ```bash tools/download-models.sh # downloads all LLM presets + whisper/VAD/TTS tools/download-models.sh smollm2-360m # or just one tools/download-models.sh list # show presets ``` (`tools/push-models.sh` is obsolete — models are in the APK now.) > **LFM-700M:** there is no `.litertlm` build of LFM2-700M on HF yet, and a > conversion requires the `litert-torch` generic-HF-export pipeline (heavy, not > guaranteed for the architecture). The SmolLM2 presets are the practical small, > non-thinking, Apache-2.0 alternatives. > **Gemma 4 E2B / E4B:** official `litert-community` LiteRT-LM builds of > `google/gemma-4-E2B-it` / `google/gemma-4-E4B-it` (Apache-2.0), using Google's > "Gemma-4 mobile" 2/4/8-bit (LUT) quantization. LiteRT-LM **memory-maps the > weights straight from the `.litertlm` file on disk** — the embedding tables are > mmap'd and never fully loaded into RAM — so the ~2.6 GB / ~3.7 GB files don't > need to fit in memory. Both support up to 32k context. Android benchmarks > (S26 Ultra, 2048 ctx): E2B ~1.7 GB CPU / ~0.7 GB GPU process RAM at > ~47–52 tok/s; E4B ~3.3 GB CPU / ~0.7 GB GPU at ~18–22 tok/s — E4B really wants > a 12 GB-class phone **and** the GPU backend. Settings → *Context window* > (LiteRT-LM `maxNumTokens`) caps the KV cache so bigger models fit in less RAM. > **Gemma 4 12B** (`gemma-4-12B-it.litertlm`, ~6.5 GB) is the largest option — > realistically it needs a 16 GB-class device with the GPU backend. > **TTS voices:** only the bundled `en_US-lessac-medium` voice is used. In-app > voice downloads were removed; Settings → *Voice (VITS)* just shows the > preinstalled voice (blank the model path for text-only chat). > **Whisper (STT):** the bundled model is `ggml-tiny.en-q4_0.bin` (~45 MB). > Better variants are available as Settings → *Whisper model* presets > (auto-downloaded on selection): unquantized tiny.en / tiny (FP16, ~75 MB), > tiny.en Q8_0 (~42 MB), or base.en Q8_0 (~78 MB). Q4_0 builds are **not** > published on Hugging Face (`tools/download-models.sh` quantizes them locally), > so Q8_0 is the closest official base.en quantized download. Larger models are > slower but more accurate, especially for non-English speech (use the > multilingual `tiny`). ## Running & testing 1. Grant the **Record audio** permission when prompted — listening starts automatically, so there is no Start button. 2. Say something. You should see `Listening → Transcribing → Thinking → Speaking`, and hear the reply sentence-by-sentence. 3. **Barge-in:** start talking mid-reply. Playback stops at once and your question is picked up as the next turn. 4. **Remote LLM:** open Settings → *Remote (OpenAI API)* → set base URL (e.g. `http://:8000/v1`) + model name, save & restart. Works with llama.cpp server, vLLM, Ollama, etc. **Context length:** Settings → *Context length* (512…32768). The app sends it as `n_ctx`, which llama.cpp honours per-request. "Server default" sends nothing — the server decides (vLLM/Ollama manage context server-side). 5. **Stop** cancels TTS + generation and re-arms listening. 6. **Visualizer mode:** tap the eye icon in the top bar to hide the transcript. The screen then shows only a minimal audio-level bar visualizer and the sentence that is currently being spoken, rendered in a cursive glowing font. 7. **Mini screen + submit photo:** tap the PiP icon to enable a small draggable floating screen (grant the *Display over other apps* permission) that shows the currently-spoken sentence in the glowing cursive font, with a distinct submit-photo button (opens the system photo picker, like the chat's photo icon). While the mini screen is active a **foreground service with a persistent "listening" notification** keeps the microphone running in the background (Android 14+). Image input only works with a **Gemma 4** model on the local LiteRT-LM backend; anything else shows an error. 8. **Sampling:** Settings → *Sampling (local)* (temperature / top-p / top-k for LiteRT-LM) and *Sampling (remote)* (temperature / top-p / max tokens for the OpenAI-compatible endpoint). Tips: - Use headphones or a quiet room: without acoustic echo cancellation support the assistant may "barge in" on itself. The capture uses `VOICE_COMMUNICATION` + `AcousticEchoCanceler` when available. - TTS can be disabled by blanking the TTS model path in Settings (text-only chat) for quick LLM testing. - GPU backend: for the **local** LiteRT-LM backend this is LiteRT's GPU delegate ("ML Drift", OpenCL on Android) — roughly doubles decode speed and cuts process RAM a lot (Gemma E4B: ~3.3 GB → ~0.7 GB). For the **remote** backend the GPU is on the server (llama.cpp CUDA/Metal/OpenCL, vLLM, …), not in the app. Enable/disable it in Settings → *GPU backend*. ## Project layout ``` app/src/main/java/com/example/voicechatbot/ MainActivity.kt Compose UI + settings dialog config/AppConfig.kt persisted settings & model paths pipeline/VoicePipelineEngine.kt state machine, barge-in, sentence streaming audio/AudioRecorder.kt 16 kHz capture (+AEC) audio/TtsPlayer.kt AudioTrack queue with generation-based barge-in vad/VadEngine.kt Silero VAD (sherpa-onnx) stt/WhisperTranscriber.kt whisper.cpp JNI wrapper stt/WhisperNative.kt JNI declarations llm/ChatClient.kt backend interface llm/LitertChatClient.kt LiteRT-LM local inference (Flow streaming) llm/OpenAiChatClient.kt OpenAI-compatible remote client (SSE) llm/SentenceSplitter.kt token stream → sentence boundaries llm/PromptBuilder.kt system prompt tts/TtsEngine.kt VITS/Piper synthesis (sherpa-onnx) app/src/main/cpp/ whisper.cpp CMake + JNI bridge third_party/whisper.cpp vendored (committed, v1.7.5) tools/download-models.sh model fetch + Q4 quantize (local) tools/push-models.sh adb push models to device ``` ## Notes & caveats - **LiteRT-LM** requires a high-end device (Pixel 8 / S23 class). Engine init takes a few seconds and is done off the UI thread. - On interrupt, the current conversation context is dropped (clean slate for the next turn). - This is a test app: no wake word, no streaming partial transcription, no foreground service (screen must stay on), single ABI (`arm64-v8a`).