Reachy_OpenWebUI / VOICE_MODE_PLAN.md
Jacid23's picture
Mark Phase 3 first cut shipped (v0.6.6.0)
e19b4a8
|
Raw
History Blame Contribute Delete
13.6 kB
# Voice Mode Plan β€” make the bot hear like OpenWebUI's browser voice mode
Status: PLAN (nothing here is implemented yet except where marked DONE)
Repo: `T:\reachy-mini\HF_Repos\Reachy_OpenWebUI_v2` β†’ Space `Jacid23/Reachy_OpenWebUI`
Bots: `.203` reachy-wireless (desk), `.204` reachy-mini (Pi 5)
Rule: NO remote daemon/app restarts β€” the user reboots the bot himself.
After any app update from the dashboard: check camera frames; if dead run
`/venvs/apps_venv/bin/pip install -U reachy-mini==1.9.0` (user or assistant-with-permission) β€” see memory notes.
---
## 1. The riddle, answered
Browser voice mode, Conduit, and the bot all use the SAME OpenWebUI backend:
batch Whisper STT (`/api/v1/audio/transcriptions`), streaming chat, sentence-level
TTS (`/api/v1/audio/speech`). **None of them are realtime models.** They are all
cascades. The browser and Conduit feel realtime because of three things the bot
lacked β€” all fixable:
1. **A working VAD with proper hysteresis.** Conduit uses Silero VAD (the `vad`
Dart package) with a dual-threshold envelope. The bot's Silero ONNX port was
broken from day one (missing the v5 64-sample context β€” returned prob ~0.001
on real speech) so it silently ran on a crude RMS energy gate. FIXED in
v0.6.5.7 (2026-07-19), verified on-bot: speech now scores mean 0.86.
2. **A clean audio front-end.** Phone/browser mics get OS-level AEC, noise
suppression, and AGC for free. The bot's mic array is BETTER hardware β€” an
XMOS-class DSP with on-chip AEC, beamforming, and AGC β€” but it is barely
configured (only `PP_AGCGAIN=5.0` is applied).
3. **Tuned endpointing + latency masking.** Conduit ships proven constants;
the bot's were guesses made to compensate for the broken VAD.
Conclusion: there is no architectural reason the bot can't feel like browser
voice mode. Torch is NOT required β€” the fixed ONNX VAD is now provably correct.
## 2. Evidence collected (2026-07-19)
### Conduit's proven VAD envelope
From `T:\reachy-mini\openwebui_related\conduit\lib\features\chat\services\voice_input_service.dart`:
| Constant | Value | Meaning |
|---|---|---|
| sample rate / frame | 16000 / 512 | same as bot |
| positiveSpeechThreshold | **0.6** | prob to ENTER speech |
| negativeSpeechThreshold | **0.35** | prob to STAY in speech (hysteresis!) |
| preSpeechPadFrames | 16 (~512 ms) | pre-roll kept before trigger |
| minSpeechFrames | 8 (~256 ms) | discard shorter blips |
| endSpeechPadFrames | 6 (~192 ms) | audio kept after end |
| redemptionFrames | 4..max (~128 ms+) | silence tolerated before ending |
Controller: `.../voice_mode/chat_voice_mode_controller.dart` (1650 lines) β€”
pauses mic during assistant speech (no barge-in), batch STT, sentence TTS.
### Bot hardware audio DSP (huge, untapped)
Daemon exposes runtime audio-DSP control:
- `POST http://<bot>:8000/api/audio/config/apply` (ApplyAudioConfigRequest)
- `GET http://<bot>:8000/api/audio/config/parameter/{name}`
Parameter families found in
`/venvs/mini_daemon/.../reachy_mini/media/audio_control_utils.py`:
`AEC_*` (echo canceller: AECSILENCELEVEL, FILTER_LENGTH, PATHCHANGE, …),
`AEC_FIXEDBEAMS*` (beamforming: azimuth/elevation/gating), `PP_AGCGAIN`, HPF etc.
Currently only `PP_AGCGAIN=5.0` is applied at app start.
### Current bot pipeline state (post today's fixes)
- Silero ONNX VAD fixed (context window) β€” Space v0.6.5.7.
- Single VAD threshold now 0.3, RMS fallback raised to 0.10 (was falsely
triggering at 0.014), mic capture `Headset,0` back at 37/60 (62%).
- Echo handling = VAD fully suppressed while TTS plays
(`_suppress_vad_until` / `_active_pipeline_count` in
`src/Reachy_OpenWebUI/sub_apps/conversation_app/local/handler.py` step 5)
β†’ no barge-in, and trailing suppression windows eat the user's next utterance.
- STT observed slow: 4563 ms for 9 s audio on the GB10 (needs investigation β€”
browser voice hits the same endpoint; short utterances + warm model are fast).
- Known infra gotchas: SDK-downgrade-on-update, media-stream wedge on remote
restart (hence the no-restart rule), pipewire masked on both bots.
## 3. Gap analysis
| Piece | Browser/Conduit | Bot today | Gap |
|---|---|---|---|
| VAD model | Silero, working | Silero ONNX, working since v0.6.5.7 | none |
| VAD envelope | dual threshold + pads | single threshold + chunk counts | Phase 1 |
| Mic front-end | OS AEC/NS/AGC | XMOS AEC/beamform barely configured | Phase 2 |
| Echo/barge-in | mic paused during TTS | VAD hard-suppressed during TTS | Phase 3 (can EXCEED them) |
| STT | same backend | same backend, seen slow | Phase 4 |
| Latency masking | call UI feedback | none | Phase 4 |
| Transport | none (local mic) | daemon media stream (fragile) | out of scope; reboot rule |
## 4. The plan
### Phase 1 β€” Adopt Conduit's VAD envelope β€” **DONE, shipped v0.6.5.8**
Implemented 2026-07-19: dual-threshold hysteresis in vad.py (enter=threshold,
stay=threshold-0.25 floor 0.15; RMS gate diagnostic-only when model loaded);
defaults now threshold 0.6, onset 3, silence-end 16 (~512ms), min-speech 8.
On-bot verified: speech 127/149 chunks / 3 segments; 4s noise 0 triggers.
NOTE: user's persisted vad_threshold=0.3 overrides the new default β€” after
updating the app, POST /vad {"vad_threshold": 0.6} (or set in settings UI).
### (original Phase 1 text follows)
File: `src/Reachy_OpenWebUI/sub_apps/conversation_app/vad.py` + `local/handler.py`.
1. Add dual-threshold hysteresis to `SileroVAD.is_speech`: enter at
`threshold_enter` (default 0.6), remain while prob β‰₯ `threshold_stay`
(default 0.35). Keep RMS fallback only as a diagnostic (log when it WOULD
have fired; do not let it trigger).
2. Map envelope to Conduit values in handler chunking (512 @16 kHz frames):
pre-roll (`lookback_buffer`) β‰₯ 16 frames, min speech 8 frames, end pad 6,
silence-end (redemption) configurable 4–30 frames (expose in settings as
today's `vad_*_chunks`).
3. Keep everything settable via the existing `/vad` endpoint + settings UI;
change only the defaults.
4. Bump version, push. User updates bots (then SDK-check ritual).
Acceptance: VAD probe shows silence ≀0.05 prob, speech β‰₯0.6; no clipped first
syllables; utterance ends within ~0.5 s of stopping; zero triggers from room
noise at normal mic gain (62%).
### Phase 2 β€” Wake the XMOS front-end β€” **CORE DONE, shipped v0.6.5.9**
2026-07-19: root cause was the app's own inherited startup config
(`audio/startup_config.py`): MIN_NS/NN 0.8 (suppressor passing 80% of noise)
+ AGC max gain 10x. Now MIN_NS/NN 0.15, AGC gain/max 4.0 β€” user-verified live:
noise floor collapsed to faint fan hum. HPF already at mode 2 (max).
Remaining Phase 2 items: beamforming (AEC_FIXEDBEAMS*) experiments, verify
persistence after user's reboot+update, STT latency (11s for 9s clip on GB10 β€”
CPU whisper; fix is OpenWebUI-side GPU STT server or Deepgram).
User's hand-tuned analog mixer sweet spot (card 0, alsamixer, 2026-07-19):
`Headset,0` β‰ˆ βˆ’8 dB, `Headset,1` β‰ˆ βˆ’22 dB. Persist with `sudo alsactl store`.
Caution: the app's mic-volume slider overwrites Headset,0.
### (original Phase 2 text follows)
Read current values first with `GET /api/audio/config/parameter/{name}`.
1. Inventory: dump all `AEC_*` and `PP_*` current values into a file for
baseline (script it; read-only).
2. Experiment matrix (apply one at a time via `POST /api/audio/config/apply`,
verify with the VAD probe logs + recorded STT quality):
- `AEC_HPFONOFF` on (kill low-frequency rumble/fan)
- `PP_AGCGAIN` sweep (current 5.0; try 2–8) with capture fixed at 62%
- `AEC_FIXEDBEAMSONOFF` + azimuth toward the user's usual position
- AEC on with far-end reference (see Phase 3)
3. Whatever wins: persist by having the app apply it at startup (the app
already applies `PP_AGCGAIN` β€” extend that config block; find it via
`grep -rn PP_AGCGAIN` in the app/daemon startup path) or via daemon config.
Caution: apply-parameter is live but survives ↔ reboot semantics unknown β€”
verify persistence after the user's next reboot before relying on it.
Acceptance: speech probs at conversation distance β‰₯0.8; noise floor rms < 0.01
at 62% gain; STT word error noticeably down (subjective A/B is fine).
### Phase 3 β€” Barge-in β€” **FIRST CUT SHIPPED v0.6.6.0 (untested on hardware)**
2026-07-19: handler no longer discards mic audio during playback. VAD keeps
running on the live stream; ~320ms sustained speech (BARGE_IN_CHUNKS=10 at
0.6 enter threshold) calls _interrupt_current_response and listening resumes
instantly (lookback keeps the interjection's start). Env knobs:
REACHY_BARGE_IN=0 reverts to half-duplex; REACHY_BARGE_IN_CHUNKS tunes
sensitivity. RISK: if the XMOS AEC leaks our own voice, the bot may
interrupt itself β€” raise chunks or disable, then investigate AEC far-end
(AEC_AECCONVERGED probe while TTS plays).
Test protocol after update+reboot: (1) normal turn works; (2) talk over the
bot mid-reply β†’ it stops and handles the interjection; (3) stay silent
through a long reply β†’ it must NOT self-interrupt.
### (original Phase 3 text follows)
Today the app deafens itself while speaking (handler step 5). With the XMOS
AEC cancelling the bot's own speaker from the mic signal, we can listen while
talking β€” which neither browser voice mode nor Conduit does.
1. Verify AEC has a far-end reference on this hardware (speaker loopback):
check `AEC_NUM_FARENDS`, `AEC_FAR_MIC_INDEX`, `AEC_AECCONVERGED` while TTS
plays (read-only probes while user runs a conversation).
2. If converged AEC is real: replace the hard `_suppress_vad_until` window
with "VAD active during playback, but require prob β‰₯ enter-threshold for N
consecutive frames (e.g. 10) to interrupt"; on trigger, call the existing
interrupt path (`_interrupt_current_response`) then treat as new utterance.
3. If AEC is not usable: fall back to half-duplex but shorten the trailing
suppression (currently eats speech after TTS ends) to ≀300 ms.
Acceptance: user can talk over the bot and it stops and listens (like
ChatGPT realtime); no self-triggering from its own voice.
### Phase 4 β€” Latency: measure, then mask
CORRECTED 2026-07-19: NO network problem β€” the running apps on both bots
already use the direct LAN route http://172.30.30.15:3001 (9.9ms), set via
the app UI (UI value overrides the .env OPENWEBUI_URL fallback; do not
diagnose from the .env file β€” ask the running app: GET :7860/status β†’
openwebui_url). The .env fallbacks were aligned to :3001 anyway. Other
routes that exist: openweb.sunrisecablema.com β†’ 172.30.30.200 (LAN reverse
proxy) and gb10.sunrisecablema.com (Cloudflare-proxied, WAN) β€” for external
clients, not the bots. Network is NOT a latency lever; whisper speed is.
1. Investigate the 4.5 s STT: time `POST /api/v1/audio/transcriptions` from
the bot with a 3 s wav (script exists in session history). If slow on GB10,
check OpenWebUI's STT engine setting (faster-whisper model size / GPU use)
β€” same win applies to every client.
2. Use the ported LatencyTracker (`REACHY_LATENCY_DETAIL=1`) to get per-stage
numbers for 10 turns; attack the biggest stage only.
3. Masking: on `Speech detected` end (STT upload start), fire an instant
"listening/thinking" cue β€” antenna twitch or attitude accent move via the
existing reactions/moves layer. Cheap and transforms perceived latency.
4. Optional eager-STT: at speech-END, we already upload immediately; consider
ALSO uploading a partial at 2 s into long utterances so Whisper is warm
(needs care with OpenWebUI; low priority).
### Phase 5 (optional, the real realtime) β€” streaming ASR transplant
Port lyon_chatbox cascade's streaming ASR provider layer
(`T:\reachy-mini\HF_Repos\lyon_chatbox\src\lyon_chatbox\cascade\asr\` β€”
`base_streaming.py`, `deepgram.py`) into conversation_app as an alternative
STT path: audio streams continuously, Deepgram does server-side endpointing
and partials, local VAD leaves the critical path entirely. Partials also feed
the transcript-reactions layer for mid-sentence robot reactions.
Cost: Deepgram API key + cloud dependency for STT only (LLM/TTS stay local).
This is the only phase that changes the architecture; Phases 1–4 likely make
it unnecessary.
## 5. Order of work & effort
1. Phase 1 β€” one sitting, app-only, ships via Space update.
2. Phase 2 β€” experiments on `.203` with user present (mic A/B), an afternoon.
3. Phase 4.1/4.3 β€” quick wins alongside Phase 2.
4. Phase 3 β€” after Phase 2 confirms AEC; the flagship feature.
5. Phase 5 β€” only if still unsatisfied.
## 5b. Torch on the bots (if ever needed β€” user-verified command)
Not required by any current phase (ONNX VAD works). If a phase needs torch
(torch-native Silero, GLiNER entity reactions, torchaudio):
`uv pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cpu`
(run inside /venvs/apps_venv; CPU wheels, aarch64-compatible.)
## 6. Hard-won constraints (do not relearn these)
- Do NOT restart daemon/app remotely; wedges media streams. User reboots.
- Dashboard app updates may downgrade `reachy-mini` to 1.8.0 β†’ camera dies β†’
pip-upgrade to 1.9.0 + user restarts.
- pipewire is masked on both bots β€” leave it masked.
- `_HTML_TAG_RE` in handler preserves `<|...|>` TTS tags (Higgs) β€” keep it.
- Mic capture control is `Headset,0` on card 0 (0–60 scale); the app's
mic_volume writes it. 37 (62%) is the sane baseline.
- Journald is volatile on the bots β€” capture evidence before reboots.