| --- |
| language: |
| - en |
| license: other |
| tags: |
| - voice |
| - text-to-speech |
| - speech-to-text |
| - real-time-audio |
| - dialectic-training |
| - zagent-evaluation |
| pretty_name: Zymatica Voice LLM |
| --- |
| |
| # Zymatica Voice LLM (ZymaticaVoice) |
| ### π Powered by [zymatica.space](https://zymatica.space) |
|
|
| > [!NOTE] |
| > **Technical Documents:** |
| > * π **[Download Technical Whitepaper PDF](https://huggingface.co/TheAiCollectiveART/Zymatica-Voice-LLM/resolve/main/Zymatica_Voice_LLM_Whitepaper.pdf)** |
| > * π **[Read Markdown Whitepaper](zymatica_voice_llm_whitepaper.md)** |
|
|
|  |
|
|
| **Zymatica Voice LLM** is an ultra-low-latency real-time voice call communication link designed to connect edge clients with large language models using advanced network compression. The system allows hands-free, microphone-based vocal calls with rapid verbal replies, mimicking natural human-to-human telephone interactions. |
|
|
| --- |
|
|
| ## π The Invention & Architecture |
|
|
| Traditional voice systems suffer from high latency due to serialized text-to-speech (TTS) and automatic speech recognition (ASR) pipelines, combined with large audio payload transfer times. ZymaticaVoice solves this through three core architectural breakthroughs: |
|
|
| ### 1. Sumerian Level 9 Audio Compression |
| By moving away from heavy Base64 string transmission (which introduces a 33% data size bloat), the server compresses raw 16-bit PCM WAV audio buffers using **Level 9 zlib deflate compression** (the maximum compression density). |
| * **Results:** Reduces HTTP network payloads by **50% to 75%**, dramatically accelerating delivery times over thin-pipe channels. |
| * **On-the-Fly Decoding:** The web client decompresses the binary buffer instantly in memory using the browser's native `DecompressionStream("deflate")` API before routing it directly to the browser's audio buffer, keeping memory footprint minimal. |
|
|
| ### 2. Sentence-Splitting & Double-Buffered Pre-fetching |
| Instead of waiting for the LLM to complete a paragraph before starting voice synthesis, ZymaticaVoice uses a pipeline split structure: |
| 1. The backend parses responses on sentence boundaries. |
| 2. The web page fetches and plays the first sentence immediately. |
| 3. While the user is listening to sentence $i$, a background thread asynchronously pre-fetches, downloads, and decompresses sentence $i+1$. |
| 4. **Transition Lag:** The player transitions between segments with exactly **0ms gap**. |
|
|
| ### 3. Latency-Hiding Routing |
| * **Fast LLM completions** are routed through high-throughput endpoints (Groq Llama 3.1 8B at `>400 tokens/sec`, Nvidia NIM, or OpenAI `gpt-4o-mini`). |
| * **Continuous browser-native transcription (ASR)** transcribes user audio as they speak, delivering final text with **0ms lag** as soon as the user stops talking. |
|
|
| --- |
|
|
| ## π Licenses Attribution Chart |
|
|
| We acknowledge and thank the creators of the open-source libraries that make the standalone pipeline run. Refer to the [LICENSE](LICENSE) file for complete details. |
|
|
| | Component Name | Author / Maintainer | Primary License | Description | |
| |------------------|-----------------------|-----------------|--------------------------------------------------| |
| | **Sumerian Level 9 Deflate** | zymatica.space | zymatica.space License | Maximum zlib deflate audio compression & browser decompression pipeline | |
| | **Double-Buffered Pre-fetch** | zymatica.space | zymatica.space License | Sentence-split pre-fetching audio playback queue | |
| | **Z Agent ZRDT Loop** | zymatica.space | zymatica.space License | Simulated dialectic dialogue & dual-observer reinforcement training loop | |
| | **Zymatica Voice Auditor** | zymatica.space | zymatica.space License | Standard audit logs, host environment signature, and MD5 cryptographic trace framework | |
| | **Language-U Cognitive Route** | zymatica.space | zymatica.space License | Sub-150ms prompt routing & key redundancy layer | |
| | **PHSS Steering Hooks** | zymatica.space | zymatica.space License | Transformer layer hooks for hidden-state vector steering | |
| | **Cuneiform-U v3 Range Coder** | zymatica.space | zymatica.space License | 6D semantic coordinate classification & adaptive arithmetic range coding engine | |
| | **Dialectic Memory System** | zymatica.space | zymatica.space License | Two-pass LLM memory extraction, Cuneiform-U seed backup, and generative decompression | |
| | **Self-Recursive Calibrator** | zymatica.space | zymatica.space License | Closed-loop prediction calibration with LLM-generated prompt patching | |
| | **Brand Assets & Logo** | TheAiCollective.art | TheAiCollective.art license | Official Zymatica brand names, visual logos, and artworks | |
| | VibeVoice | Microsoft | MIT License | Optional local 7B TTS model generation codebase | |
| | edge-tts | rany2 | MIT License | Lightweight Microsoft Edge TTS wrapper engine | |
| | aiohttp | Aio-libs team | Apache 2.0 | Asynchronous HTTP server and client framework | |
| | soundfile | Bastian Bechtold | BSD 3-Clause | Audio file writing utilities | |
| | PyTorch | Meta AI | BSD-style | Backend tensor computation library | |
| | NumPy | NumPy Developers | BSD 3-Clause | Multi-dimensional array handling | |
| | SciPy | SciPy Developers | BSD 3-Clause | Signal processing and Fourier transforms | |
| | transformers | Hugging Face | Apache 2.0 | Deep learning model configurations and loaders | |
| | safetensors | Hugging Face | Apache 2.0 | Lossless weight serialization formats | |
| | ChromaDB | Chroma | Apache 2.0 | Vector database for semantic embedding storage | |
|
|
| --- |
|
|
| ## π Usage & Documentation |
|
|
| | Document | Description | |
| |---|---| |
| | [Whitepaper (PDF)](https://huggingface.co/TheAiCollectiveART/Zymatica-Voice-LLM/resolve/main/Zymatica_Voice_LLM_Whitepaper.pdf) | Full technical whitepaper | |
| | [Whitepaper (Markdown)](zymatica_voice_llm_whitepaper.md) | Markdown version with all 11 sections | |
| | [Compression Protocol](COMPRESSION_PROTOCOL.md) | 9-level compression architecture documentation | |
| | [Compression Benchmark](benchmark_compression_protocol.py) | Runnable benchmark across all compression layers | |
|
|
| --- |
|
|
| ## π οΈ Setup & Installation Instructions |
|
|
| ### Prerequisites |
| * Python 3.9+ |
| * Active API keys for one or more fast completion providers: |
| - **Groq API Key** (highly recommended for `>400 tok/s` response times) |
| - **NVIDIA NIM API Key** |
| - **OpenAI API Key** |
|
|
| ### 1. Clone & Install Dependencies |
| Install dependencies from `requirements.txt`: |
| ```bash |
| pip install -r requirements.txt |
| ``` |
|
|
| ### 2. Configure Environment Variables |
| Create a `.env` file in the root directory and add your keys: |
| ```env |
| # Fast LLM Providers (At least one is required) |
| GROQ_API_KEY=your_groq_api_key_here |
| NVIDIA_API_KEY=your_nvidia_api_key_here |
| OPENAI_API_KEY=your_openai_api_key_here |
| ``` |
|
|
| ### 3. Run the Voice Server |
| Launch the application: |
| ```bash |
| python app.py --host 0.0.0.0 --port 5000 |
| ``` |
|
|
| --- |
|
|
| ## π± Demo Instructions |
|
|
| 1. Open your web browser and navigate to `http://localhost:5000`. |
| 2. Allow microphone access when prompted by the browser. |
| 3. Click the **Establish Comm-Link** button. You will hear an activation beep tone. |
| 4. Speak into your microphone. When you stop speaking: |
| * The page immediately logs your transcription in the CRT console. |
| * Zymatica's responses are generated, split, compressed, and streamed. |
| * The visualizer canvas displays live audio waveforms. |
| 5. Click **Terminate Link** or press `Escape` to close the call connection. |
|
|
| --- |
|
|
| ## π‘οΈ Error Handling Mechanisms |
|
|
| ZymaticaVoice includes built-in safeguards to ensure continuous call stability: |
| * **LLM key redundancy:** The server queries Groq first. If Groq fails or is unconfigured, it attempts Nvidia NIM, followed by OpenAI. If all API integrations fail, it serves a local static voice template to prevent call drops. |
| * **ASR failure protection:** If the browser doesn't support the native Web Speech API (e.g. Firefox/Safari configuration limits), it falls back gracefully to standard form text fallback in the console log. |
| * **Microphone blockage detection:** If a microphone permission is rejected or blocked, a local warning is output on the CRT console and microphone UI buttons change to warn the user without crashing the thread loop. |
|
|