Spaces:
Paused
Paused
File size: 4,732 Bytes
bc7d00b 578cfab bc7d00b 578cfab bc7d00b 578cfab bc7d00b 578cfab bc7d00b 578cfab | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 | ---
title: Gemma Avatar
emoji: π£οΈ
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
short_description: Talk to Gemma 4 face to face, with a 3D lip-synced avatar
models:
- google/gemma-4-31B-it
- nvidia/parakeet-tdt-1.1b
- Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
---
# Gemma Avatar
Realtime voice chat with a 3D talking-head avatar. Same AI stack as the
[smolagents/hf-realtime-voice](https://huggingface.co/spaces/smolagents/hf-realtime-voice)
Space, but the orb visualization is replaced by a [TalkingHead](https://github.com/met4citizen/TalkingHead)
3D avatar with real-time audio-driven lip-sync.
> **Rebuild note**: this Space was reconstructed after the original repo was
> lost. The frontend (avatar, worklets, WebSocket client, HeadAudio lip-sync)
> is a faithful recovery, vendored from `smolagents/hf-realtime-voice` and
> `met4citizen/HeadAudio`. What's genuinely new: the avatar model itself
> (generated by [MeshForge](https://huggingface.co/spaces/Daankular/meshforge-avatar)
> from a real photo, not the original's Ready Player Me `brunette.glb`), and
> the backend β this rebuild ships **direct-connect only** (no
> load-balancer session proxy), since that needs backend credentials that
> weren't recoverable. Paste a running `ws://.../v1/realtime` endpoint into
> Settings to talk to it.
## The pipeline
```
you speak β silero-VAD β parakeet-tdt-1.1b (STT) β gemma-4-31B-it on Cerebras β Qwen3-TTS β avatar speaks
```
Transport is the OpenAI Realtime GA protocol over WebSocket against an s2s
(speech-to-speech) backend: mic PCM16 @ 16 kHz goes up as
`input_audio_buffer.append`, TTS PCM16 @ 16 kHz comes back as
`response.output_audio.delta`, transcripts stream alongside.
## How the avatar works
- **Rendering / body language** β [TalkingHead](https://github.com/met4citizen/TalkingHead)
(three.js). Blinking, breathing, idle sway, moods, hand gestures, and emoji
expressions are its built-in animation system.
- **Lip-sync** β the backend sends raw PCM only (no word timings, no visemes),
so the mouth is driven from the audio itself with
[HeadAudio](https://github.com/met4citizen/HeadAudio): an AudioWorklet that
classifies MFCC frames into Oculus visemes (~50 ms latency, fully in-browser).
The s2s playback worklet is routed into TalkingHead's audio graph
(`audioAnalyzerNode β audioSpeechGainNode β reverb β speakers`) and HeadAudio
taps the speech gain node.
- **The model plays the avatar** β three function tools are declared to the
backend: `set_mood`, `make_hand_gesture`, `make_facial_expression`. Gemma
calls them mid-conversation (smiles when greeting, shrugs when unsure,
thumbs-up when agreeing).
- **Choreography** β client statuses drive presence: the avatar makes eye
contact when you start talking, gestures with its hands on new utterances,
and barge-in clears the playback buffer so the mouth settles instantly.
## Run it
```bash
bun install
bun run dev
```
Open http://localhost:7860, tap the settings icon, paste a direct
`ws://.../v1/realtime` endpoint URL, then tap **Start talking**.
`?fakemic=1` starts a session with a silent synthetic mic (useful for testing
the full loop without a microphone β trigger a reply from the console with
`getClient().requestResponse()`).
## Layout
```
index.ts Bun server: HTML import + static assets + /api/config
index.html App shell (avatar hero, caption, subtitles, settings)
src/app.js Session wiring, tool executor, UI state
src/avatar.js AvatarStage: TalkingHead + HeadAudio + choreography
src/s2s/s2s-ws-client.js Realtime WS client (vendored from smolagents/hf-realtime-voice;
orb removed, injectable output node + worklet base URL + shared-ctx close)
src/s2s/codec.js PCM/base64 + transcript helpers (vendored, unchanged)
src/vendor/headaudio.min.mjs HeadAudio node class (vendored, unchanged)
public/worklets/ mic-capture + audio-playback AudioWorklets (vendored, unchanged)
public/vendor/ HeadAudio worklet processor + viseme model (vendored, unchanged)
public/avatars/meshforge.glb Default avatar (generated from a photo by MeshForge)
```
## Notes
- The avatar GLB must have a Mixamo-compatible rig plus ARKit and Oculus-viseme
blend shapes β see [Daankular/meshforge-avatar](https://huggingface.co/spaces/Daankular/meshforge-avatar)
for how this one was generated.
- TalkingHead owns the AudioContext; the s2s client is handed `head.audioCtx`
and never closes it.
- Everything animation-related runs on requestAnimationFrame β a backgrounded
tab freezes the avatar (audio keeps playing).
|