vast-voice-stack / README.md
mohameddalii's picture
Upload folder using huggingface_hub
b0666d9 verified
|
Raw
History Blame Contribute Delete
3.69 kB
---
license: apache-2.0
tags:
- voice
- stt
- tts
- sglang
- vllm
- arabic
- blackwell
- vast
---
# Dual-5090 voice stack (STT + LLM + TTS)
Bootstrap scripts that bring up a known-good Arabic voice stack on **2× RTX 5090**
(Vast.ai / similar). Tuned for Blackwell (sm_120): no FP8 KV / ngram on the LLM,
Moss TTS forced to SDPA, CUDA graphs off, HF cache on `/dev/shm` when disk is tight.
| Service | Port | GPU | Model | Runtime |
|---------|------|-----|-------|---------|
| LLM | `8000` | 0 (shared w/ STT) | `ibm-granite/granite-4.1-8b` | SGLang |
| STT | `8002` | 0 (shared w/ LLM) | `CohereLabs/cohere-transcribe-arabic-07-2026` | vLLM 0.19 |
| TTS | `8003` | 1 (dedicated) | `Rabe3/Moss-Saudi-3` | sglang-omni |
## One-time setup on a new machine
```bash
# 1) HF token (write/read for this private repo; also needed for gated STT)
export HF_TOKEN=hf_...
# 2) Clone this private repo (replace USER with your HF username)
git clone https://USER:$HF_TOKEN@huggingface.co/USER/vast-voice-stack ~/agent-deploy
cd ~/agent-deploy
chmod +x *.sh
# 3) Disk: need ~40G+ free for two CUDA envs, OR leave HF cache on /dev/shm
# (default in common.sh). Prefer a host volume / larger disk when possible.
# 4) Launch (sequential install — do not parallelize first boot)
./deploy_all.sh
```
Upload from a machine that already has the scripts:
```bash
export HF_TOKEN=hf_... # token for the target account
./push_to_hf.sh # creates private USER/vast-voice-stack and uploads
```
Health checks: `curl -sf localhost:8000/health` / `:8002` / `:8003`.
Restart clean:
```bash
./deploy_all.sh restart-clean
```
## Load test (Locust)
```bash
./run_locust_ui.sh --local-only # or without flag to expose via Vast portal port 10100
```
Custom metrics: `llm_ttfb`, `tts_ttfa`, `stt_ttfb`, `pipe_e2e`. On this layout TTS is
the concurrency ceiling (~5–6 RPS full-WAV), not the LLM.
## What these scripts already fix
- **Disk**: sequential env install, `PIP_NO_CACHE_DIR`, shared `~/sglang-env` for LLM+TTS, HF cache → `/dev/shm/hf_home` when overlay is ~32G.
- **sglang-omni**: not on PyPI — clone + `uv pip install -e` with protobuf override.
- **STT**: transformers from git for `cohere_asr`; prefer local HF snapshot if present.
- **GPU1 budget**: TTS dedicated at `MEM_FRACTION=0.72` (colocated Moss OOM’d at 0.85).
- **GPU0 budget**: STT `MEM_UTIL=0.18` + LLM `MEM_FRACTION=0.68`.
- **LLM on Blackwell**: bf16 KV, no ngram speculative decoding (FP8/ngram caused device asserts).
- **Moss on Blackwell**: patch FA→SDPA in vocoder + audio tokenizer; `cuda_graph: false`.
Load test:
```bash
./run_loadtest.sh # TTS + pipeline concurrency ramp (stream+pcm)
./run_loadtest.sh pipe # pipeline only
./run_locust_ui.sh # interactive UI
```
## Env knobs
| Variable | Default | Meaning |
|----------|---------|---------|
| `HF_TOKEN` | — | Required for gated STT download |
| `MODEL` | per-script | Override model id or local path |
| `PORT` | 8000/8002/8003 | Bind port |
| `MEM_FRACTION` / `MEM_UTIL` | see scripts | GPU memory share |
| `LOG_DIR` | `~/agent-logs` | Server logs |
## Layout
```
common.sh # toolchain, HF home, venvs, omni + Blackwell patches
start_stt.sh / start_llm.sh / start_tts.sh
deploy_all.sh
moss_tts_local_blackwell.yaml # Moss pipeline config for 5090
locustfile.py / run_locust_ui.sh
```
## Notes
- Models are **not** vendored here (~30G). Scripts pull them via Hugging Face Hub.
- `/dev/shm` HF cache is **lost on reboot** — re-download or point `HF_HOME` at a persistent volume.
- TTS returns full WAV (not streaming); latency grows under load even when LLM TTFT stays low.