File size: 11,008 Bytes
878a301
 
 
 
 
 
9603b3c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
> The .litertlm models on this card already use the QAT that is discussed in the blog post. The most popular file, `gemma-4-E2B-it.litertlm`, uses a mixture of int2, int4 and int8 to keep it small, fast and efficient.

https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/discussions/30#6a2c2e1ac539f50158631020



# Voice Chatbot (Android) β€” interruptible, on-device voice chat pipeline

A Kotlin/Compose Android app implementing a fully local, **interruptible**
voice chatbot:

```
mic ─► Silero VAD (sherpa-onnx)
        β”‚  endpointing + barge-in
        β–Ό
   Whisper tiny.en Q4 (whisper.cpp)   ──►  transcript
        β–Ό
   LLM (streaming):
        β”œβ”€ Local:  LiteRT-LM (`.litertlm`, e.g. Qwen3-0.6B)
        └─ Remote: OpenAI-compatible `/v1/chat/completions` (SSE)
        β–Ό
   SentenceSplitter ──► per-sentence VITS TTS (sherpa-onnx + Piper voice)
        β–Ό
   AudioTrack (playback; barge-in stops it instantly)
```

**Barge-in:** speak while the assistant is talking β†’ playback and generation
stop immediately and your new utterance becomes the next turn.

## Components

| Piece | Runtime | Model | Size |
|---|---|---|---|
| STT | whisper.cpp (JNI, built via NDK/CMake) | `ggml-tiny.en-q4_0.bin` (Q4) | ~45 MB |
| VAD | sherpa-onnx AAR (official prebuilt `sherpa-onnx-1.13.4.aar`) | `silero_vad.onnx` | ~2 MB |
| TTS | sherpa-onnx VITS (Piper) | `vits-piper-en_US-lessac-medium` (+ `espeak-ng-data`) | ~62 MB |
| LLM (local) | LiteRT-LM `com.google.ai.edge.litertlm:litertlm-android:0.15.0` | `litert-community/Qwen3-0.6B` β†’ `Qwen3-0.6B.litertlm` (dynamic INT8) | 586 MB |
| LLM (remote) | OkHttp + kotlinx.serialization (OpenAI-compatible chat completions) | any remote model, e.g. `qwen2.5-0.5b-instruct` | n/a |

All models are Apache-2.0 or MIT except the Piper voice/GPL espeak-ng data
(bundled in the sherpa-onnx tts-model tarball) β€” fine for local testing.

## Prerequisites

- macOS/Linux with Android SDK (platform 35, build-tools 36, NDK `27.2.12479018`, CMake 3.22.1),
  JDK 17, and a connected Android device (arm64, 8 GB RAM recommended).
- `adb` on PATH.

## Build

Toolchain: **AGP 8.13.2 Β· Gradle 8.13 (wrapper) Β· Kotlin 2.3.21 Β· JDK 17**.
whisper.cpp is vendored under `third_party/` (committed), so the project builds
out-of-the-box once the SDK/NDK are present.

```bash
# 0. JDK 17 (Gradle 8.13 does not support newer JDKs)
#    macOS/Homebrew: brew install openjdk@17
export JAVA_HOME=/opt/homebrew/opt/openjdk@17/libexec/openjdk.jdk/Contents/Home

# 0b. Point Gradle at the Android SDK (or set ANDROID_HOME)
echo "sdk.dir=$HOME/Library/Android/sdk" > local.properties

# 1. (first time) install NDK/cmake if missing
"$ANDROID_HOME"/cmdline-tools/latest/bin/sdkmanager \
    "platforms;android-35" "ndk;27.2.12479018" "cmake;3.22.1"

# 2. Build (Gradle downloads the sherpa-onnx AAR into app/libs on first run)
./gradlew :app:assembleDebug

# 3. Install
adb install -r app/build/outputs/apk/debug/app-debug.apk
```

Unit tests: `./gradlew :app:testDebugUnitTest`

> Only if `third_party/whisper.cpp` is ever missing (e.g. a fresh export that
> dropped it): `tools/fetch-deps.sh` re-clones it.

## Models

**Models are bundled in the APK** (`assets/models/`) and copied to the app's
private internal storage on first launch β€” no adb push needed. Storage used on
the device: `~240 MB` under `/data/data/com.example.voicechatbot/files/models/`.

> The APK is ~500 MB because of the bundled models; devices need **~1.5 GB free**
> to install it.

**Preinstalled:** SmolLM2 135M (LLM) + Whisper tiny.en Q4 + Silero VAD + Piper TTS.

**Download in-app:** bigger LLMs (SmolLM2 360M, Qwen3 0.6B, Gemma 4 E2B/E4B/12B)
and Whisper models are selected in Settings via **preset chips that
auto-download the file if it isn't already saved** (progress shown on the chip).


| id | Model | Size | Preinstalled | Thinking | License |
|---|---|---|---|---|---|
| `smollm2-135m` | SmolLM2 135M | 136 MB | βœ… | no | Apache-2.0 |
| `smollm2-360m` | SmolLM2 360M | 356 MB | download | no | Apache-2.0 |
| `qwen3-0.6b` | Qwen3 0.6B int8 | 586 MB | download | yes (stripped) | Apache-2.0 |
| `gemma-4-e2b` | Gemma 4 E2B (mobile 2/4/8-bit) | 2583 MB | download | yes (stripped) | Apache-2.0 |
| `gemma-4-e4b` | Gemma 4 E4B (mobile 2/4/8-bit) | 3654 MB | download | yes (stripped) | Apache-2.0 |
| `gemma-4-12b` | Gemma 4 12B | 6548 MB | download | yes (stripped) | Apache-2.0 |

To rebuild with bundled models, the files must be present in `./models/`:

```bash
tools/download-models.sh          # downloads all LLM presets + whisper/VAD/TTS
tools/download-models.sh smollm2-360m   # or just one
tools/download-models.sh list           # show presets
```

(`tools/push-models.sh` is obsolete β€” models are in the APK now.)

> **LFM-700M:** there is no `.litertlm` build of LFM2-700M on HF yet, and a
> conversion requires the `litert-torch` generic-HF-export pipeline (heavy, not
> guaranteed for the architecture). The SmolLM2 presets are the practical small,
> non-thinking, Apache-2.0 alternatives.

> **Gemma 4 E2B / E4B:** official `litert-community` LiteRT-LM builds of
> `google/gemma-4-E2B-it` / `google/gemma-4-E4B-it` (Apache-2.0), using Google's
> "Gemma-4 mobile" 2/4/8-bit (LUT) quantization. LiteRT-LM **memory-maps the
> weights straight from the `.litertlm` file on disk** β€” the embedding tables are
> mmap'd and never fully loaded into RAM β€” so the ~2.6 GB / ~3.7 GB files don't
> need to fit in memory. Both support up to 32k context. Android benchmarks
> (S26 Ultra, 2048 ctx): E2B ~1.7 GB CPU / ~0.7 GB GPU process RAM at
> ~47–52 tok/s; E4B ~3.3 GB CPU / ~0.7 GB GPU at ~18–22 tok/s β€” E4B really wants
> a 12 GB-class phone **and** the GPU backend. Settings β†’ *Context window*
> (LiteRT-LM `maxNumTokens`) caps the KV cache so bigger models fit in less RAM.
> **Gemma 4 12B** (`gemma-4-12B-it.litertlm`, ~6.5 GB) is the largest option β€”
> realistically it needs a 16 GB-class device with the GPU backend.

> **TTS voices:** only the bundled `en_US-lessac-medium` voice is used. In-app
> voice downloads were removed; Settings β†’ *Voice (VITS)* just shows the
> preinstalled voice (blank the model path for text-only chat).

> **Whisper (STT):** the bundled model is `ggml-tiny.en-q4_0.bin` (~45 MB).
> Better variants are available as Settings β†’ *Whisper model* presets
> (auto-downloaded on selection): unquantized tiny.en / tiny (FP16, ~75 MB),
> tiny.en Q8_0 (~42 MB), or base.en Q8_0 (~78 MB). Q4_0 builds are **not**
> published on Hugging Face (`tools/download-models.sh` quantizes them locally),
> so Q8_0 is the closest official base.en quantized download. Larger models are
> slower but more accurate, especially for non-English speech (use the
> multilingual `tiny`).

## Running & testing

1. Grant the **Record audio** permission when prompted β€” listening starts
   automatically, so there is no Start button.
2. Say something. You should see `Listening β†’ Transcribing β†’ Thinking β†’
   Speaking`, and hear the reply sentence-by-sentence.
3. **Barge-in:** start talking mid-reply. Playback stops at once and your
   question is picked up as the next turn.
4. **Remote LLM:** open Settings β†’ *Remote (OpenAI API)* β†’ set base URL
   (e.g. `http://<your-server>:8000/v1`) + model name, save & restart.
   Works with llama.cpp server, vLLM, Ollama, etc.
   **Context length:** Settings β†’ *Context length* (512…32768). The app sends
   it as `n_ctx`, which llama.cpp honours per-request. "Server default" sends
   nothing β€” the server decides (vLLM/Ollama manage context server-side).
5. **Stop** cancels TTS + generation and re-arms listening.
6. **Visualizer mode:** tap the eye icon in the top bar to hide the transcript.
   The screen then shows only a minimal audio-level bar visualizer and the
   sentence that is currently being spoken, rendered in a cursive glowing font.
7. **Mini screen + submit photo:** tap the PiP icon to enable a small draggable
   floating screen (grant the *Display over other apps* permission) that shows
   the currently-spoken sentence in the glowing cursive font, with a distinct
   submit-photo button (opens the system photo picker, like the chat's photo
   icon). While the mini screen is active a **foreground service with a
   persistent "listening" notification** keeps the microphone running in the
   background (Android 14+). Image input only works with a **Gemma 4** model on
   the local LiteRT-LM backend; anything else shows an error.
8. **Sampling:** Settings β†’ *Sampling (local)* (temperature / top-p / top-k for
   LiteRT-LM) and *Sampling (remote)* (temperature / top-p / max tokens for the
   OpenAI-compatible endpoint).

Tips:

- Use headphones or a quiet room: without acoustic echo cancellation support
  the assistant may "barge in" on itself. The capture uses
  `VOICE_COMMUNICATION` + `AcousticEchoCanceler` when available.
- TTS can be disabled by blanking the TTS model path in Settings (text-only
  chat) for quick LLM testing.
- GPU backend: for the **local** LiteRT-LM backend this is LiteRT's GPU
  delegate ("ML Drift", OpenCL on Android) β€” roughly doubles decode speed and
  cuts process RAM a lot (Gemma E4B: ~3.3 GB β†’ ~0.7 GB). For the **remote**
  backend the GPU is on the server (llama.cpp CUDA/Metal/OpenCL, vLLM, …), not
  in the app. Enable/disable it in Settings β†’ *GPU backend*.

## Project layout

```
app/src/main/java/com/example/voicechatbot/
  MainActivity.kt                Compose UI + settings dialog
  config/AppConfig.kt            persisted settings & model paths
  pipeline/VoicePipelineEngine.kt  state machine, barge-in, sentence streaming
  audio/AudioRecorder.kt         16 kHz capture (+AEC)
  audio/TtsPlayer.kt             AudioTrack queue with generation-based barge-in
  vad/VadEngine.kt               Silero VAD (sherpa-onnx)
  stt/WhisperTranscriber.kt      whisper.cpp JNI wrapper
  stt/WhisperNative.kt           JNI declarations
  llm/ChatClient.kt              backend interface
  llm/LitertChatClient.kt        LiteRT-LM local inference (Flow streaming)
  llm/OpenAiChatClient.kt        OpenAI-compatible remote client (SSE)
  llm/SentenceSplitter.kt        token stream β†’ sentence boundaries
  llm/PromptBuilder.kt           system prompt
  tts/TtsEngine.kt               VITS/Piper synthesis (sherpa-onnx)
app/src/main/cpp/                 whisper.cpp CMake + JNI bridge
third_party/whisper.cpp           vendored (committed, v1.7.5)
tools/download-models.sh          model fetch + Q4 quantize (local)
tools/push-models.sh              adb push models to device
```

## Notes & caveats

- **LiteRT-LM** requires a high-end device (Pixel 8 / S23 class). Engine init
  takes a few seconds and is done off the UI thread.
- On interrupt, the current conversation context is dropped (clean slate for
  the next turn).
- This is a test app: no wake word, no streaming partial transcription, no
  foreground service (screen must stay on), single ABI (`arm64-v8a`).