gepard / README.md
akhaliq's picture
akhaliq HF Staff
Move nemo-toolkit install to runtime to unblock gradio 6.20 build
98ed9bf
|
Raw
History Blame Contribute Delete
2.47 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: GEPARD
emoji: πŸ†
colorFrom: yellow
colorTo: red
sdk: gradio
sdk_version: 6.20.0
python_version: '3.12'
app_file: app.py
pinned: false
license: apache-2.0
short_description: Spotted text-to-speech with voice cloning and text-CFG

πŸ† GEPARD TTS

Inference Space for the Gepard autoregressive speech model (nineninesix/gepard-1.0 β€” a Qwen3.5 backbone with 32 FSQ audio heads, a Q-Former voice-cloning compressor, and DPO post-training).

Features

  • Preset speakers β€” bundled voices stored as pre-encoded codec tokens (speakers/*.pt), no codec pass needed at request time.
  • Voice cloning β€” record with the microphone or upload a clip; the reference is encoded with the NeMo nano codec and compressed into 8 speaker prefix tokens.
  • Text classifier-free guidance β€” cfg_scale/cfg_frames sharpen text and voice adherence (defaults: temperature=0.3, cfg_scale=3).
  • All generation knobs are exposed under Generation settings.

Architecture

The Space runs on gradio.Server β€” a FastAPI app with Gradio's queueing engine on top. The custom dark-themed UI in index.html is served at / via @app.get("/"); the synthesis pipeline is exposed at @app.api("/synthesize"), so requests flow through Gradio's queue (concurrency control, ZeroGPU allocation, gradio_client compatibility) while the frontend stays a self-contained HTML/CSS/JS bundle.

Configuration

config.yaml selects the checkpoint, codec, preset speakers and default generation parameters β€” re-point the Space without touching code.

Notes

  • The model repo is private: set the HF_TOKEN Space secret.
  • ZeroGPU: the model is loaded once at startup; each request only runs generation inside the GPU context.
  • create_env.py orchestrates a 4-step install at app startup β€” it must run before any ML import in app.py:
    1. Pin huggingface-hub>=1.2,<2.0 (compatible with both gradio 6.20 and nemo-toolkit).
    2. pip install --no-deps nemo-toolkit[tts]==2.4.0 (avoiding a downgrade of hub by the resolver).
    3. Force-reinstall transformers==5.3.0 (Qwen3.5 backbone that the Gepard checkpoint was trained on).
    4. Cap numpy<2.0 so the codec/NeMo stack stays on numpy 1.x.
  • nemo-toolkit is installed at runtime (not in requirements.txt) to keep gradio 6.20's huggingface-hub>=1.2 constraint resolvable at build time β€” NeMo's transformers<=4.52 would otherwise pull hub<1.0.