File size: 8,859 Bytes
c0d1edb 2a8dd3f c0d1edb 2a8dd3f c0d1edb 2a8dd3f c0d1edb 2a8dd3f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 | ---
title: HF Realtime Voice
emoji: ποΈ
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
short_description: Voice chat over WebSocket against a HF speech-to-speech
hf_oauth: true
---
# Minimal Conversation App (S2S backend, **WebSocket** transport)
Drop-in alternative to [`amir-tfrere/minimal-conversation-app-s2s-backend`](https://huggingface.co/spaces/amir-tfrere/minimal-conversation-app-s2s-backend)
that uses the **WebSocket** route of the Hugging Face speech-to-speech
backend instead of the WebRTC SDP proxy. Same load balancer, same
`/session` handshake, same UI, same orb. Just a different wire.
## How it works
1. App POSTs `<lb_url>/session` (empty JSON body).
2. The LB picks a ready compute (round-robin) and returns:
```json
{
"session_id": "...",
"websocket_url": "wss://<compute>/v1/realtime",
"connect_url": "wss://<compute>/v1/realtime?session_token=<JWT>",
"session_token": "<JWT>",
"pending_timeout_s": 60
}
```
3. App opens a WebSocket **directly** on `connect_url` (no rewrite to
`https://`; unlike the WebRTC client which POSTs an SDP offer).
4. Server pushes `session.created` on connect. Client replies with
`session.update` (OpenAI Realtime **GA** schema: `session.audio.input`,
`session.audio.output`, `session.output_modalities`).
5. Client streams mic audio as PCM16 16 kHz mono base64 chunks
(`input_audio_buffer.append`, one frame every ~40 ms).
6. Server pushes `response.output_audio.delta` (PCM16 24 kHz mono base64)
and transcript deltas.
The backend exposes one concurrent session per compute (same as WebRTC
mode); the LB pins the session via a signed `session_token`.
## Why WebSocket instead of WebRTC
| | WebRTC (original) | WebSocket (this) |
|---|---|---|
| Transport | UDP + Opus 48 kHz + ICE/STUN | TCP + raw PCM16 |
| NAT traversal | needs STUN, can fail on corporate / cellular | none, works everywhere TCP is allowed |
| Audio quality | excellent (Opus, jitter buffer, FEC) | good (raw PCM, simple ring buffer) |
| Latency | lowest (~50-150 ms) | low (~150-300 ms typical) |
| Echo cancellation | browser AEC active on the WebRTC track | browser AEC active via `getUserMedia` constraints |
| Debuggability | needs `chrome://webrtc-internals` | `wscat` / DevTools network tab |
| Mobile data | sometimes blocked (UDP) | always works (HTTPS+WSS) |
## Backend requirement
This app talks to the WebSocket route `@app.websocket("/v1/realtime")`
defined in
[`websocket_router.py`](https://github.com/huggingface/speech-to-speech/blob/feat/webrtc-transport/src/speech_to_speech/api/openai_realtime/websocket_router.py)
on the **`feat/webrtc-transport`** branch. The same compute serves both
the WebRTC POST and the WebSocket upgrade on the same path; no backend
change required.
Smoke-test from the shell:
```bash
LB="https://kaa1l6rplzb1gg3y.us-east-1.aws.endpoints.huggingface.cloud"
curl -X POST "$LB/session" -H "Content-Type: application/json" -d '{}'
# -> { "connect_url": "wss://<compute>/v1/realtime?session_token=..." }
# Feed connect_url into a wscat / websocat and you should get a
# session.created event back immediately.
```
## Tools
The assistant can call two tools mid-conversation (toggle them from the **Tools**
button, top-right):
- **Web search** β Google results via Serper.dev, proxied server-side so the key
never reaches the browser. Set `SERPER_API_KEY` as a Space secret. Without it,
the tool is disabled unless the user pastes their own key in the Tools panel.
- **Camera** β while enabled, a live self-view shows bottom-left; when the model
calls the tool, the current frame is sent to the vision-language model so it can
see what you're showing it.
## Connecting to a backend
The app connects **directly** to a speech-to-speech server's realtime WebSocket β
no load balancer, no `/session` step. Set the URL in two ways:
- **`LOAD_BALANCER_URL` env** (served via `/api/config`) provides the default URL
shown in Settings β handy for the deployed Space.
- **Settings β Speech-to-speech server URL** lets you override it: paste a full
`connect_url` (`wss://host/v1/realtime?...`) or a bare host like `localhost:8080`
(the app adds `/v1/realtime`).
**Settings β Restart** reconnects with the current voice, instructions and URL.
## Usage limits
Conversation time is metered per UTC day by sign-in tier (see `limiter.py` /
`auth.py`), but **only on the deployed Space** β metering turns on only when BOTH
`LOAD_BALANCER_URL` and `SPACE_ID` (injected automatically by the HF Space
runtime) are present. Running locally β even with `LOAD_BALANCER_URL` exported β
leaves the app unmetered. Tunable via env:
| Env | Default | What |
|-----|---------|------|
| `LIMIT_ANON_SEC` | `300` | Daily seconds for anonymous visitors (5 min) |
| `LIMIT_FREE_SEC` | `600` | Daily seconds for signed-in non-PRO users (10 min) |
| `UNLIMITED_ORGS` | _(adds to defaults)_ | Extra HF org names whose members get **unlimited** usage, like PRO |
| `USAGE_HASH_SECRET` | _(random)_ | HMAC secret for hashing identity keys + signing the anon cookie |
PRO members are always unlimited. Members of `cerebras`, `HuggingFaceM4`,
`smolagents`, and `pollen-robotics` are unlimited out of the box (shown as
"Team", not "PRO"); set `UNLIMITED_ORGS=my-team` to add more. Matched
case-insensitively against the user's organisations from HF OAuth.
## Run locally
The app is now a small FastAPI server (it serves the front-end *and* the search
proxy from one container).
```bash
pip install -r requirements.txt
export SERPER_API_KEY=... # optional; web search is disabled without it
export LOAD_BALANCER_URL=... # optional; default s2s server URL (set it in Settings otherwise)
uvicorn server:app --reload --port 7860
# or, matching production: docker build -t s2s . && docker run -p 7860:7860 -e SERPER_API_KEY=... -e LOAD_BALANCER_URL=... s2s
```
Then open <http://localhost:7860/>, click the orb, allow the mic, talk.
> Browsers require **HTTPS or `localhost`** for `getUserMedia()` (mic + camera).
> `127.0.0.1` and `localhost` both work; plain `http://192.168.x.y` does NOT.
## Settings (stored in `localStorage`)
| Key | What |
|-----|------|
| Load balancer URL | Base URL of your S2S deployment. App POSTs `<lb>/session`. |
| Voice | Qwen3-TTS speaker name (Aiden, Ryan, Dylan, Eric, Ono_Anna, Serena, Sohee, Uncle_Fu, Vivian) |
| Instructions | System prompt sent in `session.update` once the WS opens |
LocalStorage keys are namespaced `s2s.ws.*` so this app's settings do
NOT collide with the WebRTC variant.
## Files
| File | Role |
|------|------|
| `index.html` | Single page, orb + settings modal (identical UI to the WebRTC app) |
| `main.js` | State machine, settings, tools, camera, noise-gate UI wiring |
| `ui/chat.js` | `ChatView`: history panel, ephemeral bubbles, transcript/tool streaming |
| `ui/account.js` | `Account`: HF login chip + popover, daily-limit modal |
| `ui/dom.js` | Shared helpers: `$`, `escHtml`, `truncateError`, `DEBUG` |
| `auth.py` | HF OAuth + per-request identity (tier, hashed keys) |
| `limiter.py` | SQLite per-day talk-time budget (chunked server-clock reservation) |
| `ws/s2s-ws-client.js` | WebSocket handshake + OpenAI Realtime GA protocol |
| `ws/codec.js` | base64 <-> PCM helpers + transcript extraction (pure) |
| `ws/orb-visualizer.js` | `OrbVisualiser`: FFT bands -> orb CSS custom properties |
| `worklets/mic-capture.js` | AudioWorklet: 48 kHz Float32 -> 16 kHz Int16 PCM, posts ~40 ms chunks |
| `worklets/audio-playback.js` | AudioWorklet: 24 kHz Float32 ring buffer -> 48 kHz, linear interp, fade in/out |
| `style.css` | Orb animations, layout, dark theme (verbatim from the WebRTC app) |
## Audio pipeline notes
- **Input**: `getUserMedia({ echoCancellation, noiseSuppression, autoGainControl })`
feeds the `mic-capture` worklet at the `AudioContext` rate. The worklet
resamples to 16 kHz (boxcar lowpass + decimation on the 48 -> 16 fast
path, linear interpolation fallback for odd rates) and packs Int16 LE.
- **Output**: `response.output_audio.delta` decodes to Int16 -> Float32
and is posted to the `audio-playback` worklet. The worklet maintains a
per-context ring buffer, linearly interpolates 24 -> 48, and applies
short 32-frame fades on entry/exit to suppress clicks.
- **Barge-in**: when the server VAD detects user speech mid-response
(`input_audio_buffer.speech_started` while `ai-speaking`), the client
posts `{ kind: "clear" }` to the playback worklet to wipe the queue
immediately. The server itself cancels the in-flight response.
## Credits
- Backend: [huggingface/speech-to-speech](https://github.com/huggingface/speech-to-speech) on `feat/webrtc-transport`
- UI verbatim from `amir-tfrere/minimal-conversation-app-s2s-backend` (Pollen Robotics Γ Hugging Face)
|