Spaces:
Running
Running
| title: Wllama | |
| emoji: π¦ | |
| colorFrom: green | |
| colorTo: pink | |
| sdk: static | |
| pinned: true | |
| short_description: Run GGUF directly on your browser! | |
| license: mit | |
| header: mini | |
| custom_headers: | |
| cross-origin-embedder-policy: require-corp | |
| cross-origin-opener-policy: same-origin | |
| cross-origin-resource-policy: cross-origin | |
| Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference | |
| ## Running locally / offline | |
| Serve this folder with the COOP/COEP headers the Space uses (required for | |
| multithreading + the model cache), then open the printed URL: | |
| ```bash | |
| python3 serve.py # http://localhost:8000 | |
| ``` | |
| What's offline-ready after one-time downloads: | |
| - **VAD + TTS**: bundled locally (`vad/`, `piper/`, `voices/`, `onnx/`) β always work offline. | |
| - **LLM model (GGUF)**: download once from the *Manage models* page; it is cached in the browser. | |
| - **ASR (Whisper)**: use the **Voice modules β Install** button on the *Manage models* page. | |
| It downloads the whisper model and runs the ONNX runtime from the local `asr/` folder, | |
| so speech-to-text works fully offline after install. | |
| ## Voice-to-voice mode | |
| Click **π Voice** in the top bar for a full-screen voice interface: | |
| - Hands-free: speak β transcribed (whisper) β answered (LLM) β read aloud (piper TTS). | |
| - A neon audio visualizer reacts to the assistant's speech. | |
| - Only sentences that have already been spoken are shown on screen, so faster | |
| text generation stays hidden until it is actually said. | |
| - On a new conversation the LLM invents a random personality for the assistant | |
| and speaks its greeting. | |
| - Click any spoken sentence (or any assistant bubble in normal chat) to | |
| **regenerate** that reply. | |