gepard / README.md
akhaliq's picture
akhaliq HF Staff
Move nemo-toolkit install to runtime to unblock gradio 6.20 build
98ed9bf
|
Raw
History Blame Contribute Delete
2.47 kB
---
title: GEPARD
emoji: πŸ†
colorFrom: yellow
colorTo: red
sdk: gradio
sdk_version: 6.20.0
python_version: '3.12'
app_file: app.py
pinned: false
license: apache-2.0
short_description: Spotted text-to-speech with voice cloning and text-CFG
---
# πŸ† GEPARD TTS
Inference Space for the **Gepard** autoregressive speech model
(`nineninesix/gepard-1.0` β€” a Qwen3.5 backbone with 32 FSQ audio
heads, a Q-Former voice-cloning compressor, and DPO post-training).
## Features
- **Preset speakers** β€” bundled voices stored as pre-encoded codec tokens
(`speakers/*.pt`), no codec pass needed at request time.
- **Voice cloning** β€” record with the microphone or upload a clip; the
reference is encoded with the NeMo nano codec and compressed into 8 speaker
prefix tokens.
- **Text classifier-free guidance** β€” `cfg_scale`/`cfg_frames` sharpen text
and voice adherence (defaults: `temperature=0.3`, `cfg_scale=3`).
- All generation knobs are exposed under **Generation settings**.
## Architecture
The Space runs on `gradio.Server` β€” a FastAPI app with Gradio's queueing
engine on top. The custom dark-themed UI in `index.html` is served at `/`
via `@app.get("/")`; the synthesis pipeline is exposed at
`@app.api("/synthesize")`, so requests flow through Gradio's queue
(concurrency control, ZeroGPU allocation, `gradio_client` compatibility)
while the frontend stays a self-contained HTML/CSS/JS bundle.
## Configuration
`config.yaml` selects the checkpoint, codec, preset speakers and default
generation parameters β€” re-point the Space without touching code.
## Notes
- The model repo is private: set the `HF_TOKEN` Space secret.
- ZeroGPU: the model is loaded once at startup; each request only runs
generation inside the GPU context.
- `create_env.py` orchestrates a 4-step install at app startup β€” it must
run before any ML import in `app.py`:
1. Pin `huggingface-hub>=1.2,<2.0` (compatible with both gradio 6.20
and `nemo-toolkit`).
2. `pip install --no-deps nemo-toolkit[tts]==2.4.0` (avoiding a downgrade
of hub by the resolver).
3. Force-reinstall `transformers==5.3.0` (Qwen3.5 backbone that the
Gepard checkpoint was trained on).
4. Cap `numpy<2.0` so the codec/NeMo stack stays on numpy 1.x.
- `nemo-toolkit` is installed at runtime (not in `requirements.txt`) to
keep gradio 6.20's `huggingface-hub>=1.2` constraint resolvable at
build time β€” NeMo's `transformers<=4.52` would otherwise pull hub<1.0.