Reachy_OpenWebUI / VOICE_MODE_PLAN.md
Jacid23's picture
Mark Phase 3 first cut shipped (v0.6.6.0)
e19b4a8
|
Raw
History Blame Contribute Delete
13.6 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade

Voice Mode Plan β€” make the bot hear like OpenWebUI's browser voice mode

Status: PLAN (nothing here is implemented yet except where marked DONE) Repo: T:\reachy-mini\HF_Repos\Reachy_OpenWebUI_v2 β†’ Space Jacid23/Reachy_OpenWebUI Bots: .203 reachy-wireless (desk), .204 reachy-mini (Pi 5) Rule: NO remote daemon/app restarts β€” the user reboots the bot himself. After any app update from the dashboard: check camera frames; if dead run /venvs/apps_venv/bin/pip install -U reachy-mini==1.9.0 (user or assistant-with-permission) β€” see memory notes.


1. The riddle, answered

Browser voice mode, Conduit, and the bot all use the SAME OpenWebUI backend: batch Whisper STT (/api/v1/audio/transcriptions), streaming chat, sentence-level TTS (/api/v1/audio/speech). None of them are realtime models. They are all cascades. The browser and Conduit feel realtime because of three things the bot lacked β€” all fixable:

  1. A working VAD with proper hysteresis. Conduit uses Silero VAD (the vad Dart package) with a dual-threshold envelope. The bot's Silero ONNX port was broken from day one (missing the v5 64-sample context β€” returned prob ~0.001 on real speech) so it silently ran on a crude RMS energy gate. FIXED in v0.6.5.7 (2026-07-19), verified on-bot: speech now scores mean 0.86.
  2. A clean audio front-end. Phone/browser mics get OS-level AEC, noise suppression, and AGC for free. The bot's mic array is BETTER hardware β€” an XMOS-class DSP with on-chip AEC, beamforming, and AGC β€” but it is barely configured (only PP_AGCGAIN=5.0 is applied).
  3. Tuned endpointing + latency masking. Conduit ships proven constants; the bot's were guesses made to compensate for the broken VAD.

Conclusion: there is no architectural reason the bot can't feel like browser voice mode. Torch is NOT required β€” the fixed ONNX VAD is now provably correct.

2. Evidence collected (2026-07-19)

Conduit's proven VAD envelope

From T:\reachy-mini\openwebui_related\conduit\lib\features\chat\services\voice_input_service.dart:

Constant Value Meaning
sample rate / frame 16000 / 512 same as bot
positiveSpeechThreshold 0.6 prob to ENTER speech
negativeSpeechThreshold 0.35 prob to STAY in speech (hysteresis!)
preSpeechPadFrames 16 (~512 ms) pre-roll kept before trigger
minSpeechFrames 8 (~256 ms) discard shorter blips
endSpeechPadFrames 6 (~192 ms) audio kept after end
redemptionFrames 4..max (~128 ms+) silence tolerated before ending

Controller: .../voice_mode/chat_voice_mode_controller.dart (1650 lines) β€” pauses mic during assistant speech (no barge-in), batch STT, sentence TTS.

Bot hardware audio DSP (huge, untapped)

Daemon exposes runtime audio-DSP control:

  • POST http://<bot>:8000/api/audio/config/apply (ApplyAudioConfigRequest)
  • GET http://<bot>:8000/api/audio/config/parameter/{name} Parameter families found in /venvs/mini_daemon/.../reachy_mini/media/audio_control_utils.py: AEC_* (echo canceller: AECSILENCELEVEL, FILTER_LENGTH, PATHCHANGE, …), AEC_FIXEDBEAMS* (beamforming: azimuth/elevation/gating), PP_AGCGAIN, HPF etc. Currently only PP_AGCGAIN=5.0 is applied at app start.

Current bot pipeline state (post today's fixes)

  • Silero ONNX VAD fixed (context window) β€” Space v0.6.5.7.
  • Single VAD threshold now 0.3, RMS fallback raised to 0.10 (was falsely triggering at 0.014), mic capture Headset,0 back at 37/60 (62%).
  • Echo handling = VAD fully suppressed while TTS plays (_suppress_vad_until / _active_pipeline_count in src/Reachy_OpenWebUI/sub_apps/conversation_app/local/handler.py step 5) β†’ no barge-in, and trailing suppression windows eat the user's next utterance.
  • STT observed slow: 4563 ms for 9 s audio on the GB10 (needs investigation β€” browser voice hits the same endpoint; short utterances + warm model are fast).
  • Known infra gotchas: SDK-downgrade-on-update, media-stream wedge on remote restart (hence the no-restart rule), pipewire masked on both bots.

3. Gap analysis

Piece Browser/Conduit Bot today Gap
VAD model Silero, working Silero ONNX, working since v0.6.5.7 none
VAD envelope dual threshold + pads single threshold + chunk counts Phase 1
Mic front-end OS AEC/NS/AGC XMOS AEC/beamform barely configured Phase 2
Echo/barge-in mic paused during TTS VAD hard-suppressed during TTS Phase 3 (can EXCEED them)
STT same backend same backend, seen slow Phase 4
Latency masking call UI feedback none Phase 4
Transport none (local mic) daemon media stream (fragile) out of scope; reboot rule

4. The plan

Phase 1 β€” Adopt Conduit's VAD envelope β€” DONE, shipped v0.6.5.8

Implemented 2026-07-19: dual-threshold hysteresis in vad.py (enter=threshold, stay=threshold-0.25 floor 0.15; RMS gate diagnostic-only when model loaded); defaults now threshold 0.6, onset 3, silence-end 16 (~512ms), min-speech 8. On-bot verified: speech 127/149 chunks / 3 segments; 4s noise 0 triggers. NOTE: user's persisted vad_threshold=0.3 overrides the new default β€” after updating the app, POST /vad {"vad_threshold": 0.6} (or set in settings UI).

(original Phase 1 text follows)

File: src/Reachy_OpenWebUI/sub_apps/conversation_app/vad.py + local/handler.py.

  1. Add dual-threshold hysteresis to SileroVAD.is_speech: enter at threshold_enter (default 0.6), remain while prob β‰₯ threshold_stay (default 0.35). Keep RMS fallback only as a diagnostic (log when it WOULD have fired; do not let it trigger).
  2. Map envelope to Conduit values in handler chunking (512 @16 kHz frames): pre-roll (lookback_buffer) β‰₯ 16 frames, min speech 8 frames, end pad 6, silence-end (redemption) configurable 4–30 frames (expose in settings as today's vad_*_chunks).
  3. Keep everything settable via the existing /vad endpoint + settings UI; change only the defaults.
  4. Bump version, push. User updates bots (then SDK-check ritual). Acceptance: VAD probe shows silence ≀0.05 prob, speech β‰₯0.6; no clipped first syllables; utterance ends within ~0.5 s of stopping; zero triggers from room noise at normal mic gain (62%).

Phase 2 β€” Wake the XMOS front-end β€” CORE DONE, shipped v0.6.5.9

2026-07-19: root cause was the app's own inherited startup config (audio/startup_config.py): MIN_NS/NN 0.8 (suppressor passing 80% of noise)

  • AGC max gain 10x. Now MIN_NS/NN 0.15, AGC gain/max 4.0 β€” user-verified live: noise floor collapsed to faint fan hum. HPF already at mode 2 (max). Remaining Phase 2 items: beamforming (AEC_FIXEDBEAMS*) experiments, verify persistence after user's reboot+update, STT latency (11s for 9s clip on GB10 β€” CPU whisper; fix is OpenWebUI-side GPU STT server or Deepgram). User's hand-tuned analog mixer sweet spot (card 0, alsamixer, 2026-07-19): Headset,0 β‰ˆ βˆ’8 dB, Headset,1 β‰ˆ βˆ’22 dB. Persist with sudo alsactl store. Caution: the app's mic-volume slider overwrites Headset,0.

(original Phase 2 text follows)

Read current values first with GET /api/audio/config/parameter/{name}.

  1. Inventory: dump all AEC_* and PP_* current values into a file for baseline (script it; read-only).
  2. Experiment matrix (apply one at a time via POST /api/audio/config/apply, verify with the VAD probe logs + recorded STT quality):
    • AEC_HPFONOFF on (kill low-frequency rumble/fan)
    • PP_AGCGAIN sweep (current 5.0; try 2–8) with capture fixed at 62%
    • AEC_FIXEDBEAMSONOFF + azimuth toward the user's usual position
    • AEC on with far-end reference (see Phase 3)
  3. Whatever wins: persist by having the app apply it at startup (the app already applies PP_AGCGAIN β€” extend that config block; find it via grep -rn PP_AGCGAIN in the app/daemon startup path) or via daemon config. Caution: apply-parameter is live but survives ↔ reboot semantics unknown β€” verify persistence after the user's next reboot before relying on it. Acceptance: speech probs at conversation distance β‰₯0.8; noise floor rms < 0.01 at 62% gain; STT word error noticeably down (subjective A/B is fine).

Phase 3 β€” Barge-in β€” FIRST CUT SHIPPED v0.6.6.0 (untested on hardware)

2026-07-19: handler no longer discards mic audio during playback. VAD keeps running on the live stream; ~320ms sustained speech (BARGE_IN_CHUNKS=10 at 0.6 enter threshold) calls _interrupt_current_response and listening resumes instantly (lookback keeps the interjection's start). Env knobs: REACHY_BARGE_IN=0 reverts to half-duplex; REACHY_BARGE_IN_CHUNKS tunes sensitivity. RISK: if the XMOS AEC leaks our own voice, the bot may interrupt itself β€” raise chunks or disable, then investigate AEC far-end (AEC_AECCONVERGED probe while TTS plays). Test protocol after update+reboot: (1) normal turn works; (2) talk over the bot mid-reply β†’ it stops and handles the interjection; (3) stay silent through a long reply β†’ it must NOT self-interrupt.

(original Phase 3 text follows)

Today the app deafens itself while speaking (handler step 5). With the XMOS AEC cancelling the bot's own speaker from the mic signal, we can listen while talking β€” which neither browser voice mode nor Conduit does.

  1. Verify AEC has a far-end reference on this hardware (speaker loopback): check AEC_NUM_FARENDS, AEC_FAR_MIC_INDEX, AEC_AECCONVERGED while TTS plays (read-only probes while user runs a conversation).
  2. If converged AEC is real: replace the hard _suppress_vad_until window with "VAD active during playback, but require prob β‰₯ enter-threshold for N consecutive frames (e.g. 10) to interrupt"; on trigger, call the existing interrupt path (_interrupt_current_response) then treat as new utterance.
  3. If AEC is not usable: fall back to half-duplex but shorten the trailing suppression (currently eats speech after TTS ends) to ≀300 ms. Acceptance: user can talk over the bot and it stops and listens (like ChatGPT realtime); no self-triggering from its own voice.

Phase 4 β€” Latency: measure, then mask

CORRECTED 2026-07-19: NO network problem β€” the running apps on both bots already use the direct LAN route http://172.30.30.15:3001 (9.9ms), set via the app UI (UI value overrides the .env OPENWEBUI_URL fallback; do not diagnose from the .env file β€” ask the running app: GET :7860/status β†’ openwebui_url). The .env fallbacks were aligned to :3001 anyway. Other routes that exist: openweb.sunrisecablema.com β†’ 172.30.30.200 (LAN reverse proxy) and gb10.sunrisecablema.com (Cloudflare-proxied, WAN) β€” for external clients, not the bots. Network is NOT a latency lever; whisper speed is.

  1. Investigate the 4.5 s STT: time POST /api/v1/audio/transcriptions from the bot with a 3 s wav (script exists in session history). If slow on GB10, check OpenWebUI's STT engine setting (faster-whisper model size / GPU use) β€” same win applies to every client.
  2. Use the ported LatencyTracker (REACHY_LATENCY_DETAIL=1) to get per-stage numbers for 10 turns; attack the biggest stage only.
  3. Masking: on Speech detected end (STT upload start), fire an instant "listening/thinking" cue β€” antenna twitch or attitude accent move via the existing reactions/moves layer. Cheap and transforms perceived latency.
  4. Optional eager-STT: at speech-END, we already upload immediately; consider ALSO uploading a partial at 2 s into long utterances so Whisper is warm (needs care with OpenWebUI; low priority).

Phase 5 (optional, the real realtime) β€” streaming ASR transplant

Port lyon_chatbox cascade's streaming ASR provider layer (T:\reachy-mini\HF_Repos\lyon_chatbox\src\lyon_chatbox\cascade\asr\ β€” base_streaming.py, deepgram.py) into conversation_app as an alternative STT path: audio streams continuously, Deepgram does server-side endpointing and partials, local VAD leaves the critical path entirely. Partials also feed the transcript-reactions layer for mid-sentence robot reactions. Cost: Deepgram API key + cloud dependency for STT only (LLM/TTS stay local). This is the only phase that changes the architecture; Phases 1–4 likely make it unnecessary.

5. Order of work & effort

  1. Phase 1 β€” one sitting, app-only, ships via Space update.
  2. Phase 2 β€” experiments on .203 with user present (mic A/B), an afternoon.
  3. Phase 4.1/4.3 β€” quick wins alongside Phase 2.
  4. Phase 3 β€” after Phase 2 confirms AEC; the flagship feature.
  5. Phase 5 β€” only if still unsatisfied.

5b. Torch on the bots (if ever needed β€” user-verified command)

Not required by any current phase (ONNX VAD works). If a phase needs torch (torch-native Silero, GLiNER entity reactions, torchaudio): uv pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cpu (run inside /venvs/apps_venv; CPU wheels, aarch64-compatible.)

6. Hard-won constraints (do not relearn these)

  • Do NOT restart daemon/app remotely; wedges media streams. User reboots.
  • Dashboard app updates may downgrade reachy-mini to 1.8.0 β†’ camera dies β†’ pip-upgrade to 1.9.0 + user restarts.
  • pipewire is masked on both bots β€” leave it masked.
  • _HTML_TAG_RE in handler preserves <|...|> TTS tags (Higgs) β€” keep it.
  • Mic capture control is Headset,0 on card 0 (0–60 scale); the app's mic_volume writes it. 37 (62%) is the sane baseline.
  • Journald is volatile on the bots β€” capture evidence before reboots.