Spaces:
Running
Running
File size: 13,564 Bytes
36da88f bb3e104 36da88f f4b0c00 d3576ea f4b0c00 36da88f e19b4a8 36da88f 750ee6b 36da88f 1583db9 36da88f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 | # Voice Mode Plan β make the bot hear like OpenWebUI's browser voice mode
Status: PLAN (nothing here is implemented yet except where marked DONE)
Repo: `T:\reachy-mini\HF_Repos\Reachy_OpenWebUI_v2` β Space `Jacid23/Reachy_OpenWebUI`
Bots: `.203` reachy-wireless (desk), `.204` reachy-mini (Pi 5)
Rule: NO remote daemon/app restarts β the user reboots the bot himself.
After any app update from the dashboard: check camera frames; if dead run
`/venvs/apps_venv/bin/pip install -U reachy-mini==1.9.0` (user or assistant-with-permission) β see memory notes.
---
## 1. The riddle, answered
Browser voice mode, Conduit, and the bot all use the SAME OpenWebUI backend:
batch Whisper STT (`/api/v1/audio/transcriptions`), streaming chat, sentence-level
TTS (`/api/v1/audio/speech`). **None of them are realtime models.** They are all
cascades. The browser and Conduit feel realtime because of three things the bot
lacked β all fixable:
1. **A working VAD with proper hysteresis.** Conduit uses Silero VAD (the `vad`
Dart package) with a dual-threshold envelope. The bot's Silero ONNX port was
broken from day one (missing the v5 64-sample context β returned prob ~0.001
on real speech) so it silently ran on a crude RMS energy gate. FIXED in
v0.6.5.7 (2026-07-19), verified on-bot: speech now scores mean 0.86.
2. **A clean audio front-end.** Phone/browser mics get OS-level AEC, noise
suppression, and AGC for free. The bot's mic array is BETTER hardware β an
XMOS-class DSP with on-chip AEC, beamforming, and AGC β but it is barely
configured (only `PP_AGCGAIN=5.0` is applied).
3. **Tuned endpointing + latency masking.** Conduit ships proven constants;
the bot's were guesses made to compensate for the broken VAD.
Conclusion: there is no architectural reason the bot can't feel like browser
voice mode. Torch is NOT required β the fixed ONNX VAD is now provably correct.
## 2. Evidence collected (2026-07-19)
### Conduit's proven VAD envelope
From `T:\reachy-mini\openwebui_related\conduit\lib\features\chat\services\voice_input_service.dart`:
| Constant | Value | Meaning |
|---|---|---|
| sample rate / frame | 16000 / 512 | same as bot |
| positiveSpeechThreshold | **0.6** | prob to ENTER speech |
| negativeSpeechThreshold | **0.35** | prob to STAY in speech (hysteresis!) |
| preSpeechPadFrames | 16 (~512 ms) | pre-roll kept before trigger |
| minSpeechFrames | 8 (~256 ms) | discard shorter blips |
| endSpeechPadFrames | 6 (~192 ms) | audio kept after end |
| redemptionFrames | 4..max (~128 ms+) | silence tolerated before ending |
Controller: `.../voice_mode/chat_voice_mode_controller.dart` (1650 lines) β
pauses mic during assistant speech (no barge-in), batch STT, sentence TTS.
### Bot hardware audio DSP (huge, untapped)
Daemon exposes runtime audio-DSP control:
- `POST http://<bot>:8000/api/audio/config/apply` (ApplyAudioConfigRequest)
- `GET http://<bot>:8000/api/audio/config/parameter/{name}`
Parameter families found in
`/venvs/mini_daemon/.../reachy_mini/media/audio_control_utils.py`:
`AEC_*` (echo canceller: AECSILENCELEVEL, FILTER_LENGTH, PATHCHANGE, β¦),
`AEC_FIXEDBEAMS*` (beamforming: azimuth/elevation/gating), `PP_AGCGAIN`, HPF etc.
Currently only `PP_AGCGAIN=5.0` is applied at app start.
### Current bot pipeline state (post today's fixes)
- Silero ONNX VAD fixed (context window) β Space v0.6.5.7.
- Single VAD threshold now 0.3, RMS fallback raised to 0.10 (was falsely
triggering at 0.014), mic capture `Headset,0` back at 37/60 (62%).
- Echo handling = VAD fully suppressed while TTS plays
(`_suppress_vad_until` / `_active_pipeline_count` in
`src/Reachy_OpenWebUI/sub_apps/conversation_app/local/handler.py` step 5)
β no barge-in, and trailing suppression windows eat the user's next utterance.
- STT observed slow: 4563 ms for 9 s audio on the GB10 (needs investigation β
browser voice hits the same endpoint; short utterances + warm model are fast).
- Known infra gotchas: SDK-downgrade-on-update, media-stream wedge on remote
restart (hence the no-restart rule), pipewire masked on both bots.
## 3. Gap analysis
| Piece | Browser/Conduit | Bot today | Gap |
|---|---|---|---|
| VAD model | Silero, working | Silero ONNX, working since v0.6.5.7 | none |
| VAD envelope | dual threshold + pads | single threshold + chunk counts | Phase 1 |
| Mic front-end | OS AEC/NS/AGC | XMOS AEC/beamform barely configured | Phase 2 |
| Echo/barge-in | mic paused during TTS | VAD hard-suppressed during TTS | Phase 3 (can EXCEED them) |
| STT | same backend | same backend, seen slow | Phase 4 |
| Latency masking | call UI feedback | none | Phase 4 |
| Transport | none (local mic) | daemon media stream (fragile) | out of scope; reboot rule |
## 4. The plan
### Phase 1 β Adopt Conduit's VAD envelope β **DONE, shipped v0.6.5.8**
Implemented 2026-07-19: dual-threshold hysteresis in vad.py (enter=threshold,
stay=threshold-0.25 floor 0.15; RMS gate diagnostic-only when model loaded);
defaults now threshold 0.6, onset 3, silence-end 16 (~512ms), min-speech 8.
On-bot verified: speech 127/149 chunks / 3 segments; 4s noise 0 triggers.
NOTE: user's persisted vad_threshold=0.3 overrides the new default β after
updating the app, POST /vad {"vad_threshold": 0.6} (or set in settings UI).
### (original Phase 1 text follows)
File: `src/Reachy_OpenWebUI/sub_apps/conversation_app/vad.py` + `local/handler.py`.
1. Add dual-threshold hysteresis to `SileroVAD.is_speech`: enter at
`threshold_enter` (default 0.6), remain while prob β₯ `threshold_stay`
(default 0.35). Keep RMS fallback only as a diagnostic (log when it WOULD
have fired; do not let it trigger).
2. Map envelope to Conduit values in handler chunking (512 @16 kHz frames):
pre-roll (`lookback_buffer`) β₯ 16 frames, min speech 8 frames, end pad 6,
silence-end (redemption) configurable 4β30 frames (expose in settings as
today's `vad_*_chunks`).
3. Keep everything settable via the existing `/vad` endpoint + settings UI;
change only the defaults.
4. Bump version, push. User updates bots (then SDK-check ritual).
Acceptance: VAD probe shows silence β€0.05 prob, speech β₯0.6; no clipped first
syllables; utterance ends within ~0.5 s of stopping; zero triggers from room
noise at normal mic gain (62%).
### Phase 2 β Wake the XMOS front-end β **CORE DONE, shipped v0.6.5.9**
2026-07-19: root cause was the app's own inherited startup config
(`audio/startup_config.py`): MIN_NS/NN 0.8 (suppressor passing 80% of noise)
+ AGC max gain 10x. Now MIN_NS/NN 0.15, AGC gain/max 4.0 β user-verified live:
noise floor collapsed to faint fan hum. HPF already at mode 2 (max).
Remaining Phase 2 items: beamforming (AEC_FIXEDBEAMS*) experiments, verify
persistence after user's reboot+update, STT latency (11s for 9s clip on GB10 β
CPU whisper; fix is OpenWebUI-side GPU STT server or Deepgram).
User's hand-tuned analog mixer sweet spot (card 0, alsamixer, 2026-07-19):
`Headset,0` β β8 dB, `Headset,1` β β22 dB. Persist with `sudo alsactl store`.
Caution: the app's mic-volume slider overwrites Headset,0.
### (original Phase 2 text follows)
Read current values first with `GET /api/audio/config/parameter/{name}`.
1. Inventory: dump all `AEC_*` and `PP_*` current values into a file for
baseline (script it; read-only).
2. Experiment matrix (apply one at a time via `POST /api/audio/config/apply`,
verify with the VAD probe logs + recorded STT quality):
- `AEC_HPFONOFF` on (kill low-frequency rumble/fan)
- `PP_AGCGAIN` sweep (current 5.0; try 2β8) with capture fixed at 62%
- `AEC_FIXEDBEAMSONOFF` + azimuth toward the user's usual position
- AEC on with far-end reference (see Phase 3)
3. Whatever wins: persist by having the app apply it at startup (the app
already applies `PP_AGCGAIN` β extend that config block; find it via
`grep -rn PP_AGCGAIN` in the app/daemon startup path) or via daemon config.
Caution: apply-parameter is live but survives β reboot semantics unknown β
verify persistence after the user's next reboot before relying on it.
Acceptance: speech probs at conversation distance β₯0.8; noise floor rms < 0.01
at 62% gain; STT word error noticeably down (subjective A/B is fine).
### Phase 3 β Barge-in β **FIRST CUT SHIPPED v0.6.6.0 (untested on hardware)**
2026-07-19: handler no longer discards mic audio during playback. VAD keeps
running on the live stream; ~320ms sustained speech (BARGE_IN_CHUNKS=10 at
0.6 enter threshold) calls _interrupt_current_response and listening resumes
instantly (lookback keeps the interjection's start). Env knobs:
REACHY_BARGE_IN=0 reverts to half-duplex; REACHY_BARGE_IN_CHUNKS tunes
sensitivity. RISK: if the XMOS AEC leaks our own voice, the bot may
interrupt itself β raise chunks or disable, then investigate AEC far-end
(AEC_AECCONVERGED probe while TTS plays).
Test protocol after update+reboot: (1) normal turn works; (2) talk over the
bot mid-reply β it stops and handles the interjection; (3) stay silent
through a long reply β it must NOT self-interrupt.
### (original Phase 3 text follows)
Today the app deafens itself while speaking (handler step 5). With the XMOS
AEC cancelling the bot's own speaker from the mic signal, we can listen while
talking β which neither browser voice mode nor Conduit does.
1. Verify AEC has a far-end reference on this hardware (speaker loopback):
check `AEC_NUM_FARENDS`, `AEC_FAR_MIC_INDEX`, `AEC_AECCONVERGED` while TTS
plays (read-only probes while user runs a conversation).
2. If converged AEC is real: replace the hard `_suppress_vad_until` window
with "VAD active during playback, but require prob β₯ enter-threshold for N
consecutive frames (e.g. 10) to interrupt"; on trigger, call the existing
interrupt path (`_interrupt_current_response`) then treat as new utterance.
3. If AEC is not usable: fall back to half-duplex but shorten the trailing
suppression (currently eats speech after TTS ends) to β€300 ms.
Acceptance: user can talk over the bot and it stops and listens (like
ChatGPT realtime); no self-triggering from its own voice.
### Phase 4 β Latency: measure, then mask
CORRECTED 2026-07-19: NO network problem β the running apps on both bots
already use the direct LAN route http://172.30.30.15:3001 (9.9ms), set via
the app UI (UI value overrides the .env OPENWEBUI_URL fallback; do not
diagnose from the .env file β ask the running app: GET :7860/status β
openwebui_url). The .env fallbacks were aligned to :3001 anyway. Other
routes that exist: openweb.sunrisecablema.com β 172.30.30.200 (LAN reverse
proxy) and gb10.sunrisecablema.com (Cloudflare-proxied, WAN) β for external
clients, not the bots. Network is NOT a latency lever; whisper speed is.
1. Investigate the 4.5 s STT: time `POST /api/v1/audio/transcriptions` from
the bot with a 3 s wav (script exists in session history). If slow on GB10,
check OpenWebUI's STT engine setting (faster-whisper model size / GPU use)
β same win applies to every client.
2. Use the ported LatencyTracker (`REACHY_LATENCY_DETAIL=1`) to get per-stage
numbers for 10 turns; attack the biggest stage only.
3. Masking: on `Speech detected` end (STT upload start), fire an instant
"listening/thinking" cue β antenna twitch or attitude accent move via the
existing reactions/moves layer. Cheap and transforms perceived latency.
4. Optional eager-STT: at speech-END, we already upload immediately; consider
ALSO uploading a partial at 2 s into long utterances so Whisper is warm
(needs care with OpenWebUI; low priority).
### Phase 5 (optional, the real realtime) β streaming ASR transplant
Port lyon_chatbox cascade's streaming ASR provider layer
(`T:\reachy-mini\HF_Repos\lyon_chatbox\src\lyon_chatbox\cascade\asr\` β
`base_streaming.py`, `deepgram.py`) into conversation_app as an alternative
STT path: audio streams continuously, Deepgram does server-side endpointing
and partials, local VAD leaves the critical path entirely. Partials also feed
the transcript-reactions layer for mid-sentence robot reactions.
Cost: Deepgram API key + cloud dependency for STT only (LLM/TTS stay local).
This is the only phase that changes the architecture; Phases 1β4 likely make
it unnecessary.
## 5. Order of work & effort
1. Phase 1 β one sitting, app-only, ships via Space update.
2. Phase 2 β experiments on `.203` with user present (mic A/B), an afternoon.
3. Phase 4.1/4.3 β quick wins alongside Phase 2.
4. Phase 3 β after Phase 2 confirms AEC; the flagship feature.
5. Phase 5 β only if still unsatisfied.
## 5b. Torch on the bots (if ever needed β user-verified command)
Not required by any current phase (ONNX VAD works). If a phase needs torch
(torch-native Silero, GLiNER entity reactions, torchaudio):
`uv pip install torch==2.5.1 torchvision==0.20.1 torchaudio==2.5.1 --index-url https://download.pytorch.org/whl/cpu`
(run inside /venvs/apps_venv; CPU wheels, aarch64-compatible.)
## 6. Hard-won constraints (do not relearn these)
- Do NOT restart daemon/app remotely; wedges media streams. User reboots.
- Dashboard app updates may downgrade `reachy-mini` to 1.8.0 β camera dies β
pip-upgrade to 1.9.0 + user restarts.
- pipewire is masked on both bots β leave it masked.
- `_HTML_TAG_RE` in handler preserves `<|...|>` TTS tags (Higgs) β keep it.
- Mic capture control is `Headset,0` on card 0 (0β60 scale); the app's
mic_volume writes it. 37 (62%) is the sane baseline.
- Journald is volatile on the bots β capture evidence before reboots.
|